Retrieval-augmented generation gets talked about as if it were a new kind of model. It isn't. RAG is a pattern: before you ask a model a question, you fetch relevant documents and place them in the prompt, so the answer is grounded in your data instead of the model's frozen memory. Useful, but not magic — and naming it helps you see where the work actually is.
It's a search problem in disguise
Most of RAG is a search problem in disguise. If retrieval surfaces the wrong passages, no amount of prompt-tuning saves you — the model will faithfully summarise the wrong thing, often with total confidence. I end up spending far more time on chunking, embeddings, and ranking than on the generation step, because that's where the quality is actually decided.
Chunking is underrated
Chunking is underrated. Split documents too small and each piece loses the context that made it meaningful; split too large and you bury the one relevant sentence in noise while burning context-window budget. There's no universal right size — it depends on your documents — which is exactly why it deserves experimentation instead of a copied-in default.
The honest failure mode
The honest failure mode is confident citation of irrelevant context. A good retrieval system doesn't just return the top results no matter what; it can decide nothing cleared the relevance bar and say "I don't have anything on that."
Top-k is not a relevance guarantee
Handing the model the top-k passages every time, regardless of quality, is how you get answers that look sourced but aren't. Gate on a similarity threshold and allow an empty result — a system that can say "nothing matched" is more trustworthy than one that always has something to cite.
Hybrid beats pure
Hybrid beats pure. Dense vector search is great for meaning — it knows "car" and "automobile" are close — but weak on exact tokens like names, error codes, and part numbers, where old-fashioned keyword search wins outright. The systems I trust run both and re-rank the merged results.
Pros
- Dense vectors match meaning — synonyms and paraphrases land close together.
- Keyword (BM25) nails exact tokens: names, error codes, SKUs, part numbers.
- Merging both and re-ranking recovers what either search alone would miss.
Cons
- Dense search alone slips on rare exact strings it never learned to embed.
- Keyword search alone is blind to paraphrase and synonymy.
- Two indexes plus a re-ranker is more moving parts to build and operate.
In practice the merged flow is a short, boring pipeline — which is exactly the point. RAG rewards search-engineering discipline far more than clever prompting:
# Hybrid retrieval: run both searches, merge, then re-rank.
dense = vector_index.search(embed(query), k=20) # matches meaning
sparse = bm25_index.search(query, k=20) # matches exact tokens
candidates = dedupe(dense + sparse)
ranked = reranker.score(query, candidates) # cross-encoder re-rank
hits = [c for c in ranked if c.score >= THRESHOLD]
return hits[:5] if hits else "I don't have anything on that."Sources & Further Reading
- 01Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al., 2020The paper that named the pattern.
- 02Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks — Reimers & Gurevych, 2019Sentence embeddings for dense retrieval.
- 03Efficient and robust approximate nearest neighbor search using HNSW graphs — Malkov & Yashunin, 2018The index most vector stores lean on.
Editorial note — An explainer on established retrieval techniques (chunking, dense vs. keyword search, re-ranking). No benchmarks or vendor-specific claims are asserted.


