RAG context compression: send fewer, better chunks to the model
RAG context compression is the step between retrieval and generation that decides which retrieved chunks are worth sending to the model, removes duplicates and off-topic passages, and fits the rest into a token budget. It is distinct from retrieval tuning, which changes what gets found, and from reranking, which changes the order. Compression changes how much the model reads, which is what you pay for.
- Top-k retrieval is optimized for recall, so it always returns more than the answer needs. The waste is structural, not a tuning mistake.
- Retrieval tuning changes what is found, reranking changes the order, compression changes how many tokens reach the model. Only the last one directly reduces input cost.
- The safest pattern is to retrieve broadly and send narrowly: raise top-k, then select, deduplicate and budget-fit before generation.
- Aggressive compression fails in predictable ways: lost citations, over-cut evidence and broken groundedness. A quality constraint on an evaluation set prevents all three.
- Measure tokens before and after alongside answer quality and groundedness. A compression step that is not measured against quality is a liability.
Why top-k retrieval sends unnecessary context
Top-k retrieval sends unnecessary context because it is designed to maximize recall, not to minimize what the model reads. A vector index returns the k nearest chunks to the query embedding, and k is chosen high enough that the right passage is almost always included. Everything else in that list is the cost of that guarantee.
The problem is not that top-k is wrong. It is that similarity is a weak proxy for usefulness. A chunk can be semantically close to the query and still add nothing: it restates a fact another chunk already covers, it discusses the topic in a different time period, or it matches on vocabulary while answering a different question. Embedding distance cannot see any of that. It only knows that the vectors are near each other.
Chunking makes it worse. Fixed-size chunking splits a document into pieces that do not respect meaning, so the passage that answers the question often straddles two chunks, and both get retrieved with their surrounding noise. Overlapping windows, which are used to fix that boundary problem, duplicate text on purpose. Every overlap is paid for twice. The noise is not free on quality either: models use evidence buried in the middle of a long context less reliably than evidence near its edges (Lost in the Middle).
The result is a retrieval stage that behaves correctly by its own metric and still produces a prompt where a large share of tokens never influence the answer. That share is what context optimization targets.
Where tokens leak in a RAG pipeline, step by step
Tokens leak at five points in a typical RAG pipeline, and each one adds to the prompt without anyone deciding it should.
- Chunking. Fixed windows with overlap create duplicate text before retrieval even runs. A 512-token chunk with 128-token overlap carries 25% repeated content at every boundary.
- Top-k. Returning 20 chunks when 3 would answer the question multiplies input tokens by roughly 6x for that request, and the extra 17 chunks are the ones least likely to matter.
- Hybrid merge. Combining dense and sparse (BM25) results improves recall, but the union of two lists is larger than either. Without deduplication, the same passage can appear from both retrievers.
- Metadata and formatting. Source names, headers, page numbers and JSON wrappers are often injected per chunk. Across 20 chunks, this framing can rival the content itself.
- Prompt assembly. System instructions, few-shot examples and history are concatenated with the retrieved context. None of it is trimmed to the current question.
Each step is reasonable in isolation. The waste comes from the fact that nothing between retrieval and generation asks the only question that matters for cost: does this chunk change the answer? That is the question a compression stage exists to answer, and it is the same question that agent pipelines face with tool output.
Retrieval tuning, reranking and compression are three different tools
Retrieval tuning changes what gets found, reranking changes the order of what was found, and compression changes how much of it reaches the model. They are often discussed as if they were interchangeable ways to improve RAG, but they act on different variables and fix different failures.
Retrieval tuning
Retrieval tuning covers chunk size, overlap, embedding model choice, top-k, hybrid search and query rewriting. Its goal is recall: make sure the answer is somewhere in the retrieved set. Lowering top-k reduces tokens, but it does so blindly, dropping results before anything has looked at them. That trades cost for recall, which is the wrong trade for a production system that has to answer correctly.
Reranking
Reranking uses a cross-encoder or a similar model that reads the query and each chunk together and produces a relevance score. It is much more accurate than embedding distance because it can see interactions between the two texts. It fixes ordering: the best chunk moves to the top. It does not, by itself, fix volume. A reranked list of 20 chunks is still 20 chunks unless something truncates it, and truncating at a fixed position is just top-k again with a better sort.
Compression and selection
Compression is the family of techniques that reduce the token count of the context that is actually sent. It includes extractive selection (keep whole chunks or sentences that pass a relevance decision), deduplication (collapse near-identical passages), budget fitting (pack the highest-value items into a fixed token limit) and abstractive summarization (rewrite content into fewer tokens). Token-level pruning tools such as Microsoft LLMLingua sit at the abstractive end. Compression is the only one of the three that directly reduces input tokens per request without lowering recall at the retrieval stage.
| Retrieval tuning | Reranking | Compression / selection | |
|---|---|---|---|
| What it optimizes | Recall: is the answer in the retrieved set | Order: is the best result first | Volume: how many tokens the model reads |
| Acts on | Index, chunking, query, k | The result list order | The result list contents and size |
| Cost per request | None at query time beyond retrieval | One cross-encoder pass per chunk | One cheap decision per chunk, optional rewrite |
| Latency | Low | Low to medium, scales with k | Low, decisions run in parallel |
| Quality risk | Low k drops the answer before anyone sees it | Low, order rarely hurts | Over-cut, lost citations if unconstrained |
| Reduces input tokens | Only by lowering k, which lowers recall | No, unless combined with a cutoff | Yes, directly |
| Use when | Recall is the problem | The right chunk is retrieved but buried | The right chunk is retrieved, but with too much company |
In practice the three are layered. Tune retrieval for recall, rerank for order, then compress for volume. Teams that skip the last step end up paying the model to do the compression stage's job at inference prices. The wider cost picture is covered in LLM cost optimization.
The pattern: retrieve broadly, send narrowly
The safest way to compress RAG context is to raise top-k so recall is not the bottleneck, then apply a cheap selection stage that decides per chunk whether it earns a place in the prompt. Recall is protected by the wide retrieval. Cost is controlled by the narrow send.
This inverts the usual instinct. Most teams lower top-k to save tokens, which is the one lever that removes the answer before it can be evaluated. Retrieving 40 chunks and sending 8 costs less than retrieving 10 and sending 10, and it is more likely to include the passage that matters, because the selection stage has more to choose from.
chunks = retrieve(query, k=40) # wide net, recall first
chunks = rerank(query, chunks) # optional: better order
kept = []
for chunk in chunks:
if not is_relevant(query, chunk): # cheap binary decision
continue
if is_duplicate(chunk, kept): # near-duplicate check
continue
kept.append(chunk)
context = fit_to_budget(kept, budget=6000) # pack by value per token
if quality_guard(query, context, chunks) == "over_cut":
context = restore_top_dropped(context, chunks)
answer = llm.generate(query, context)Three properties make this pattern safe. First, decisions are per chunk, so nothing is cut mid-passage and citations stay intact. Second, the relevance decision is cheap, so it can run on every chunk in parallel without adding meaningful latency. Third, the budget is explicit, so cost per request is bounded by design rather than by hoping the retriever returns little. Token budgets describes the budget stage in more detail.
Compression techniques, from safest to most aggressive
Compression techniques range from removing whole chunks, which preserves every surviving passage exactly, to rewriting text, which can shorten aggressively but risks changing meaning. The safer end should be exhausted before the aggressive end is used.
- Chunk-level selection. Keep or drop each chunk based on a relevance decision against the query. Surviving text is untouched, so quotes and citations remain valid. This is the largest and safest lever in most pipelines.
- Deduplication. Collapse near-identical chunks from overlapping windows, hybrid retrievers or repeated documents. Use embedding similarity or shingling with a tight threshold. Pure win when the threshold is conservative.
- Sentence-level extraction. Within a kept chunk, drop sentences that do not relate to the query. Higher savings, but a dropped sentence can remove a qualifier that changes the meaning of the one before it. Keep whole paragraphs when the domain is legal, medical or financial.
- Budget fitting. Rank surviving items by value per token and pack them into a fixed limit. This bounds cost but must never truncate mid-item. Whole items in, whole items out.
- Abstractive summarization. Rewrite long passages into shorter ones with a small model, or prune tokens with a perplexity-based method. Highest compression ratio, highest risk: the model now reads a paraphrase, not the source, and groundedness checks against the original become unreliable.
A useful rule: prefer techniques that remove items over techniques that rewrite items. Removal is auditable. You can show exactly which chunk was dropped and why. Rewriting is not, and when an answer goes wrong nobody can tell whether the source or the summary was at fault.
How aggressive compression fails, and how a quality constraint prevents it
Aggressive compression fails in three recognizable ways: it loses citations, it over-cuts the evidence the answer depends on, and it breaks groundedness by feeding the model paraphrases instead of sources. Each failure is prevented by the same mechanism, a quality constraint that the optimizer must satisfy before it is allowed to save tokens.
Lost citations
When compression cuts inside a chunk or rewrites it, the span the model would have cited no longer exists in the prompt. The answer may still be right, but it can no longer point to its source, and any citation-checking step downstream fails. Chunk-level selection avoids this entirely. Sentence-level extraction needs source offsets preserved.
Over-cut
Over-cut is when the selection stage removes a chunk that the answer needed, usually because the relevance decision was made in isolation and the chunk only mattered in combination with another. A question that requires comparing two policies needs both; a per-chunk score can rate each as weakly relevant and drop both. The fix is a guard that checks the compressed context against the question before generation and restores dropped items when coverage looks thin.
Broken groundedness
Groundedness measures whether the answer is supported by the provided context. Abstractive compression can produce a summary that is faithful in spirit but introduces a phrase the source never used. The model then repeats it, the answer looks grounded in the summary, and it is not grounded in the document. This is why groundedness must be evaluated against the original sources, not against the compressed context.
A quality constraint turns compression from a cost target into a constrained optimization: minimize tokens subject to answer quality and groundedness staying within a tolerance measured on an evaluation set. That is the design behind a quality guard. Without it, every compression setting is a guess, and the aggressive settings that look best on the cost report are the ones most likely to be quietly wrong.
An evaluation recipe for RAG context compression
Evaluate compression by running the same questions through the pipeline with and without the compression stage, then comparing tokens, answer quality and groundedness side by side. Cost numbers without quality numbers are not a result.
- Build an evaluation set. 100 to 300 real questions with reference answers and, where possible, the source passages that support them. Sample from production logs, not from the documentation's own examples.
- Freeze the retriever. Use the same index, chunking and top-k for both arms so the only variable is compression.
- Run the baseline. Record input tokens, output tokens, latency and cost per question with all retrieved chunks sent.
- Run the compressed arm. Record the same, plus the per-question removal log: which chunks were dropped and why.
- Score answer quality. Use a rubric or a judge model against the reference answer. Report the delta between arms, not the absolute score.
- Score groundedness against sources. Check that each claim in the answer is supported by a passage in the original retrieved set, not in the compressed prompt.
- Inspect the tail. Sort by quality delta and read the ten worst regressions. This is where over-cut shows up, and the removal log tells you what was dropped.
- Set the tolerance. Decide the maximum acceptable quality drop, then choose the most aggressive setting that stays inside it. Re-run monthly as the corpus changes.
The output is a report with four numbers per setting: input tokens before, input tokens after, answer quality delta, groundedness delta. That is the same shape as a savings report, and it is the only form in which a compression claim should be made.
A worked example with the token arithmetic
A support assistant that retrieves 20 chunks of 800 tokens each sends 16,000 tokens of context per question; after chunk-level selection to 4 chunks and deduplication, it sends 3,200. The figures below are illustrative, but the arithmetic is what you should reproduce on your own workload.
| Stage | Items | Tokens | Note |
|---|---|---|---|
| Retrieved (top-k = 20) | 20 chunks | 16,000 | 20 x 800, hybrid dense + BM25 |
| After deduplication | 16 chunks | 12,800 | 4 near-duplicates from overlap and retriever union |
| After relevance selection | 5 chunks | 4,000 | 11 chunks judged not to change the answer |
| After budget fit (4,000 limit) | 4 chunks | 3,200 | Lowest value-per-token chunk dropped, whole |
| Quality guard | 4 chunks | 3,200 | Coverage check passed, nothing restored |
Per request, context tokens fall from 16,000 to 3,200, a reduction of 12,800 tokens or 80%. Add a 600-token system prompt and a 200-token question to both arms and the full input goes from 16,800 to 4,000, a 76% reduction. At an illustrative input price of $3 per million tokens, that is $0.0504 before and $0.0120 after, or $0.0384 saved per request. At 50,000 requests a day, the daily saving is 50,000 x $0.0384 = $1,920, before counting the latency improvement from the model reading four chunks instead of twenty.
The number that makes this a result rather than a hope is the quality column. If the evaluation set shows answer quality within tolerance and groundedness flat or improved, which it often is when noise is removed, the saving is real. If quality dropped outside tolerance, the budget was too tight or the relevance threshold too strict, and the removal log shows which chunks to stop dropping.
This selection stage can be built in-house from a reranker score with a cutoff, a deduplication pass and a budget packer. It can also be handled by a dedicated layer such as Spendwaise, which runs the decisions, budget and guard as one call through its API and produces the before-and-after report. Either way, the pattern is the same: retrieve broadly, send narrowly, and measure both cost and quality.
Where RAG compression fits in the wider context stack
RAG context compression is one instance of a general rule: decide what the model reads before paying the model to read it. The same rule applies to tool outputs, conversation history and web results, and the same machinery handles all of them.
Retrieved chunks are usually the largest single source of prompt tokens in a RAG application, which is why compression pays off quickly there. But a production assistant rarely has only one source. Agents accumulate tool results across steps, chat products carry history into every turn, and search-augmented systems pull in whole pages. Each has its own waste pattern, described in the agent context optimization guide, and each benefits from the same select, deduplicate, budget and guard stages.
Treating context as a budgeted resource across all sources, rather than tuning each source in isolation, is the core of what practitioners now call context engineering. RAG compression is the most mature piece of it, and a good place to start because the evaluation is straightforward and the savings are visible on the first report.
Frequently asked questions
What is RAG context compression?
RAG context compression is the stage between retrieval and generation that reduces the tokens sent to the model by removing chunks that will not change the answer, collapsing duplicates and fitting the remainder into a token budget. It reduces input cost and latency without lowering retrieval recall.
Is reranking the same as compression?
No. Reranking reorders retrieved chunks by relevance but does not reduce their number or size. Compression reduces the number of tokens that reach the model. A reranker is often a useful input to compression, because its scores make a good relevance signal, but reranking alone does not lower cost.
Why not just lower top-k to save tokens?
Lowering top-k removes results before anything has evaluated them, so it trades recall for cost blindly. Retrieving a wide set and then selecting from it costs less per request and is more likely to include the passage that answers the question.
Does compressing context hurt answer quality?
It can, if it is unconstrained. Chunk-level selection with deduplication and a quality guard usually holds answer quality flat and often improves groundedness, because noise is removed. Abstractive rewriting carries more risk. The only way to know is to measure quality on an evaluation set alongside the token savings.
How much does RAG context compression save?
It depends on how much of the retrieved context is redundant or off-topic on your workload. Pipelines with high top-k, overlapping chunks and hybrid retrieval have the most waste. Measure tokens before and after on your own questions rather than relying on a published percentage.
Where does compression fail?
Three places: lost citations when text is cut mid-chunk or rewritten, over-cut when a chunk that only matters in combination with another is dropped, and broken groundedness when the model reads a paraphrase instead of the source. Each is prevented by a quality constraint evaluated against the original retrieved passages.