numpy, sentence_transformers, qdrant_client, openai, and fastapi were all absent, which is exactly the constraint the fallback path is designed for. Anything that needed a package or a key is labelled and given as an exact command.
The companion video for this topic is in the references at the bottom, and the architecture around it is in RAGs.
What a RAG pipeline is, in stages
1
Chunk
Split documents into retrieval units of roughly 250 to 400 tokens. The unit is what I retrieve, so its boundaries decide what the model can ever see.
2
Embed or index
Turn each chunk into something searchable — a vector for semantic search, an inverted index for lexical search. Both are retrieval; only one needs a model.
3
Retrieve
Score the query against the index, take top-k, and keep the provenance (document id, chunk id) attached.
4
Prompt
Assemble a context block from retrieved chunks with citation anchors, then the instruction and the question.
5
Generate
Ask the model to answer only from that context and to cite the anchors it used.
6
Evaluate
Score retrieval (recall, precision, MRR) and generation (groundedness, answer correctness) separately, because they fail differently and cost differently to fix.
The runnable version, end to end
Lexical BM25 stands in for the vector index, and a deterministic extractive selector stands in for the LLM. Everything else is a real RAG pipeline. Save asminirag.py and run python3 minirag.py.
Verified output
This is whatpython3 minirag.py printed in my environment, not a hand-written approximation:
answer-match=0.40 against answer-in-context=1.00, which I will explain in a moment because it is the most instructive number on the page.
The retrieval score, worked by hand
For the termbellamar in chunk d1#0, measured from the same run:
idf × tf-saturation. df=3 out of N=13 chunks gives idf = ln(1 + (13 − 3 + 0.5)/(3 + 0.5)) = 1.3863. The saturation term with k1=1.2, b=0.75, tf=1, chunk length 21 and avgdl=16.08 is (1 × 2.2) / (1 + 1.2 × (1 − 0.75 + 0.75 × 21/16.08)) = 0.8887. Their product, 1.2320, is that term’s contribution to the chunk’s score. Two properties I want you to see: a term appearing more times saturates toward k1 + 1 rather than growing linearly, and a chunk longer than average is penalised by b. Those are the two knobs that explain most “why is this result ranked there” conversations.
Reading the evaluation honestly
Metric arithmetic for this run, so you can check my numbers: ranks were 1, 1, 2, 1, 1, soMRR = (1 + 1 + 0.5 + 1 + 1) / 5 = 0.900 and recall@3 = 5/5 = 1.00. Per-query doc-level precision was 1/3, 1/3, 1/3, 2/3, 1/3 — the fourth query got two chunks from the same gold document inside top-3 — so precision@3 = 2.0/5 = 0.40.
Now the interesting part. Retrieval was near-perfect and generation still failed on 3 of 5 questions. I deliberately printed answer-in-context alongside answer-match to make that separation visible: the correct answer sentence was present in the retrieved context for all 5 questions, and my generator only got 2 of them.
The clearest case is how much does the airport transfer cost. Traced from the same run:
transfer does not match transfers, and cost has no lexical bridge to fee or euro — df(cost) = 0, so it contributes nothing at all. Three fixable failure modes in one query: no stemming or lemmatisation, no synonym handling, and no semantic signal. That is exactly the query-document mismatch and chunking-loses-context pair from the FDE interview syllabus, and it is why the production path below is not optional in a real system.
The production path, per stage
Each stage above has a real replacement. I did not run these in this environment because the packages and keys are absent, so treat the commands as the exact ones to execute.pip install ragas, score faithfulness (is the claim supported by context), answer relevancy, context precision, and context recall — the last two being exactly the split my table above makes with precision@3 and answer-in-context.
Order of investment in a real system, based on what actually moved metrics for me: chunking quality first (a badly bounded chunk cannot be retrieved well by any index), then hybrid retrieval with fusion, then the reranker, then the prompt, then the model swap. People buy the last one first and wonder why the answers are still wrong.
What I would change next, in order of payoff
Having the pipeline runnable means each of these is an experiment I can score rather than a guess. Chunk size. Mymax_words=25 chunks are deliberately small, and small chunks retrieve precisely but answer poorly, because the sentence that answers “what time is check-in” loses the sentence two lines later that says it applies only to loyalty members. Large chunks do the opposite: fewer vectors, cheaper index, more context per hit, and worse ranking because a chunk about five things matches a query about one of them. I sweep 128, 256, and 512 tokens and read recall@k and answer-match together — the pair always disagrees, and the disagreement is the finding.
Overlap. Zero overlap is a coin flip on boundary-spanning answers. Adding a sentence of overlap costs vectors — the 1.2x multiplier in Data size estimation — and buys back the boundary cases. It is the cheapest quality improvement available and the first thing I measure.
Top-k. With k=3 I already have recall@3 = 1.00 on this corpus, so a larger k would only add tokens and cost. On a real corpus the opposite holds: raise k for recall, then let the reranker cut it back down before the prompt. The reranker is how I get both a wide candidate pool and a small context budget.
Stemming, then semantics. The transfer versus transfers miss is fixed by stemming, which is roughly twenty lines and stdlib-free in this design. The cost versus fee miss cannot be fixed lexically at all — df(cost) = 0 means the term contributes nothing — and that is precisely the gap dense embeddings exist to close. Once a real embedding model is in place I keep BM25 alongside it and fuse with RRF, because the lexical channel is what pins exact identifiers, prices, and error codes.
Prompt and abstention. My instruction already says “if the context does not answer the question, say so”, and the offline generator implements abstention when the score is zero. That behaviour is not a nicety: in production the most expensive failure is a confident answer over an empty or wrong context, and I want an explicit, measurable abstain rate rather than a hallucination rate discovered by a customer.
Cost and latency of what I just built
The toy corpus hides the economics, so I size them. This pipeline retrieves 5 chunks of about 300 tokens, roughly 1,550 input tokens per query before instructions. Against the Why Need RAG if Gemini Can Answer alternative — putting the entire corpus in context on every turn — retrieval is not a quality decision first, it is a per-query billing decision, and the difference grows with every follow-up question in the same session.
So the pipeline has a fixed cost I amortise and a variable cost I have to budget per turn, and the tuning knobs above move the variable line directly.
Three failure modes this run demonstrates on purpose, all of them on the standard interview syllabus in What to Expect in FDE GenAI Roles: chunking losing context (boundary effects), query-document mismatch (
transfer versus transfers, cost versus fee), and stale indexes (my corpus is a dict literal; a real one must be re-embedded when the document or the model version changes). Recall versus precision is the trade I made visible by printing both.References
- Complete RAG tutorial video — the source material for this page, also listed in RAGs.
- RAGs — the pipeline as an architecture and teaching syllabus.
- Why Need RAG if Gemini Can Answer — when to skip this whole build and use a long context instead.