RAG is just search with extra steps
Retrieval-augmented generation gets written up like an architecture. It is a search query, a string concatenation and one API call — and the search query is the part that decides whether it works.
Contents
RAG — retrieval-augmented generation — has accumulated an unreasonable amount of ceremony. Diagrams with eleven boxes. Frameworks with plugin systems. Job titles.
Here is the entire technique, in one function:
async function answer(question) {
const docs = await search(question); // 1. find relevant text
const context = docs.map((d) => d.text).join('\n\n---\n\n');
return llm(`Answer using only the context below. If the answer
is not there, say you do not know.
Context:
${context}
Question: ${question}`); // 2. paste it in and ask
}
Find relevant text. Paste it into the prompt. Ask the question. That is RAG.
I am not being dismissive — this technique is genuinely how you get a model to answer questions about your company’s data without training anything, and I have shipped it more than once. But knowing that it is ninety percent a search problem changes where you spend your week, and almost everybody spends their week in the wrong place.
Why you need it at all
A language model knows what was in its training data and nothing else. It does not know your refund policy, your API’s error codes, or what your team decided in March. Ask it anyway and it will produce something plausible and wrong, because producing plausible text is precisely what it does.
You have two options. Fine-tune the model on your data — expensive, slow to update, and surprisingly bad at factual recall. Or put the facts in front of it at the moment you ask. The second one is RAG, it costs nothing to update, and when the source document changes the answer changes with it.
The part that decides everything
If the retrieval step hands over the wrong three paragraphs, no model on earth will produce the right answer. It will produce a fluent, confident answer based on the wrong three paragraphs, which is considerably worse than an error.
So: your retrieval quality is your product quality. Here is what actually moves it.
Search both ways and merge
Pure vector search misses exact identifiers. Someone searching for error code ERR_4021 gets semantic neighbours about errors in general, because the model has no idea that string is special. Pure keyword search misses paraphrase.
Run both. Merge the ranked lists. The standard merge is reciprocal rank fusion, and it is much simpler than the name:
function fuse(lists, k = 60) {
const scores = new Map();
for (const list of lists) {
list.forEach((doc, rank) => {
scores.set(doc.id, (scores.get(doc.id) ?? 0) + 1 / (k + rank + 1));
});
}
return [...scores.entries()]
.sort((a, b) => b[1] - a[1])
.map(([id]) => id);
}
A document ranked well by either method rises; a document ranked well by both wins. No tuning, no weights to guess. This one change reliably gives me the biggest quality jump in any retrieval system I have built.
Rerank the shortlist
Retrieve twenty, then have a small cross-encoder model score each one against the query properly and keep the best three. Reranking is slower per document, which is exactly why you do it on twenty and not on twenty thousand.
The economics are good: three excellent chunks cost less in tokens than ten mediocre ones and produce better answers. You save money and get better output, which is not a trade you are offered often.
Fix the question before you search it
Users do not write search queries. They write “what about the other one?” — which is meaningless on its own.
In a conversation, rewrite the question into a standalone one before retrieving. One cheap call to a small model:
Given this conversation, rewrite the final message as a standalone question. Output only the question.
This takes twenty minutes to add and fixes the single most common complaint about chat-over-docs features, which is that they work perfectly on the first question and fall apart on the second.
Make it say “I don’t know”
The instruction “if the answer is not in the context, say you do not know” works far better than people expect, and it is the difference between a feature people trust and a feature people quietly stop using.
Pair it with citations. Number the chunks, ask for the numbers back, and render them as links:
[1] refund-policy.md § Eligibility
[2] refund-policy.md § Timeframes
Citations do two jobs. They let the reader check, and they let you check — when an answer is wrong, the citation tells you instantly whether retrieval failed or the model misread. Without them, every bug report is a mystery.
When to skip it entirely
RAG is for corpora too large to send. If your entire knowledge base is 30,000 tokens, do not build a pipeline. Paste the whole thing in, cache the prefix, and go and do something more useful with the fortnight.
I have seen two teams build vector databases for document sets that fit comfortably in a single prompt. Measure the size of the thing before you architect around it.
The honest summary
Chunk your documents with their headings attached. Search by keyword and by vector, then merge. Rerank the shortlist. Rewrite follow-up questions. Demand citations. Let it say it does not know.
Six things. None of them require a framework. All of them are search engineering, which web developers have been doing for twenty-five years — this time the last step just happens to be a model instead of a results page.