AI system design · Token budgets, semantic cache
How to design a RAG pipeline
Retrieval-augmented generation (RAG) answers questions with a language model grounded in your own documents: retrieve the most relevant passages, put them in the prompt, and generate an answer with citations. It has become a standard AI system design interview question, and it is classic system design with new units - tokens instead of requests, context windows instead of memory, and cost per query instead of cost per server.
Updated · 5 min read
Requirements
- Functional: answer natural-language questions over a document corpus, with citations to the source passages.
- Functional: new and edited documents become searchable within minutes.
- Functional: users only ever see content from documents they are allowed to read.
- Non-functional: the first token of the answer appears within about two seconds; cost per query stays inside a budget; answers are grounded in the retrieved text.
Capacity estimates
| Quantity | Assumption | Result |
|---|---|---|
| Corpus | 10 million documents, ≈ 10 chunks each | ≈ 100 million chunks |
| Vector storage | 1,024-dimension float32 embeddings (4 KB each) | ≈ 400 GB; ≈ 100 GB with 8-bit quantization |
| Query load | 100 queries/s at peak | - |
| Prompt size | 8 chunks × 400 tokens + instructions + question | ≈ 4,000 input tokens per query |
| Token throughput | 4,000 in + 400 out per query | ≈ 400,000 input and 40,000 output tokens/s |
Cost per query is input tokens × input price + output tokens × output price. Because the prompt is ten times longer than the answer, the number of retrieved chunks is usually the biggest cost lever you control.
Ingestion pipeline
- A document change event goes onto a queue, so ingestion absorbs bursts and retries failures without blocking anything.
- Workers parse the document, split it into chunks of a few hundred tokens with a small overlap, and attach metadata: document ID, section, timestamp and access-control list.
- Chunks are embedded in batches and upserted into the vector index, keyed by document and chunk ID so re-ingesting an edited document replaces its old chunks.
- Keep the raw text and metadata in ordinary storage too. Changing the embedding model means re-embedding everything, and you need the source text to do it.
Query path
- Authenticate the user and resolve which documents they may read.
- Embed the question and retrieve candidates: approximate nearest-neighbour vector search, often combined with keyword search (hybrid retrieval) for names and exact terms.
- Filter candidates by permission inside the search, not after it, so restricted text never reaches the prompt.
- Rerank the top 50 candidates with a cross-encoder and keep the best 5-10.
- Assemble the prompt with numbered sources, call the model and stream tokens back to the user as they are generated.
Vector search at scale
Exact nearest-neighbour search over 100 million vectors is too slow, so vector databases use approximate indexes such as HNSW graphs or IVF partitions. They trade a little recall for orders of magnitude lower latency, and their parameters let you tune that trade-off. Shard the index by tenant or by hash when it outgrows one machine, and replicate shards for read throughput.
Where it breaks
- Latency stacks: query embedding + retrieval + reranking + generation all happen in sequence, and generation is the slowest step.
- The context window is finite: retrieve too many chunks and you exceed it or pay for tokens the model barely uses; retrieve too few and the answer misses facts.
- Bad retrieval looks like bad generation: when the right passage isn't retrieved, the model answers from general knowledge or invents details.
- Stale or leaky indexes: a slow ingestion pipeline serves outdated answers, and permission checks applied after retrieval can leak restricted text into prompts.
Cutting latency and cost
- Stream the answer, so perceived latency is time to first token rather than total generation time.
- Retrieve fewer, better chunks by reranking - the cheapest improvement to both quality and cost.
- Prompt caching: providers can cache a long, repeated prompt prefix (instructions and examples) and bill it at a reduced rate on later calls.
- Semantic cache: return a stored answer when a new question's embedding is close enough to a previous one. Scope it per tenant and permission set, and set the similarity threshold carefully - too loose and users get answers to different questions.
- Route simple questions to a smaller, cheaper model and reserve the large model for hard ones.
Evaluating quality
Keep a fixed evaluation set of real questions with known good sources. Measure retrieval separately (did the right passage appear in the top k?) from generation (is the answer faithful to the retrieved text?). Run the set on every change to chunking, embeddings, retrieval or prompts, and collect user feedback in production.
What interviewers look for
- You estimated tokens per query and connected them to latency and cost.
- You separated the ingestion path from the query path and made ingestion asynchronous.
- You enforced permissions inside retrieval.
- You reranked and streamed, and you can explain the semantic cache's correctness risk.
- You described how you would measure quality.
Frequently asked questions
What is a RAG system design interview?
+
It asks you to design a retrieval-augmented generation system at scale - for example question answering over millions of documents at a given query rate. You design ingestion, the vector index, retrieval and reranking, prompt assembly and model serving, and reason about tokens, latency, cost per query and permissions.
What is a semantic cache?
+
A cache keyed by the meaning of a query rather than its exact text. It stores past questions with their answers and returns a stored answer when a new question's embedding is similar enough. It cuts cost and latency for repetitive questions, at the risk of returning an answer to a slightly different question if the threshold is too loose.
How do you reduce the cost of a RAG system?
+
Retrieve fewer but better chunks by reranking, cache repeated prompt prefixes, add a semantic cache for repeated questions, and route simple questions to a smaller model. Cost scales with tokens, so shrinking the prompt has the largest effect.
What is the bottleneck in a RAG pipeline?
+
Usually end-to-end latency, dominated by generation, followed by retrieval quality. The context window limits how much you can retrieve, and cost per token limits how much you can afford to send.
How do you handle document permissions in RAG?
+
Store access-control metadata with every chunk and filter by the user's permissions inside the retrieval query, so restricted passages are never retrieved, never placed in a prompt and never cached across users.