Context compression with a quality check on every request.
Every reduction is a trade-off. Spendwaise checks each one after the response, learns from it, and shows you every decision it made.
A cost report can look perfect while answers get worse.
Every context reduction technique can be tuned to show a bigger saving. Cut harder and the token count falls further. The problem is that the cost dashboard improves in exactly the same way whether the removed context was noise or the one paragraph the answer depended on. Nobody notices the difference until users do.
Aggressive compression fails in three recognizable ways: it over-cuts the evidence an answer needs, it loses the passages the model would have cited, and it breaks groundedness by feeding the model paraphrases instead of sources. The RAG context compression guide covers each failure in detail. The quality check exists so that none of them can happen silently.
A check after every response, a correction before the next.
The check runs only when the relevance threshold actually removed something, and only after your response has been sent. It turns every request into feedback for the agent's next one.
The request is optimized at the agent's current threshold and the context is returned to you.
Your model gets the context. Nothing below adds a millisecond to that response.
A decision model answers one yes/no question about the kept context: is it enough to answer the question completely and accurately? The answer is a calibrated probability.
Below 50%, the agent's threshold drops by 0.15 so it keeps more. At 85% or above, it rises by 0.05 so it cuts more. In between, it stays.
The threshold stays within fixed limits, and steps down are three times larger than steps up: quality recovers faster than savings grow.
Every removed item comes with a reason.
Trust in an optimizer comes from being able to audit it. Each request records what was removed and why, so a surprising answer can be traced back to the context decision behind it.
| Reason | What it means | Typical example |
|---|---|---|
| Low relevance | The item scored below the agent's threshold for this question. | A billing policy retrieved for a shipping question. |
| Duplicate | The same information is already in the context. | Two overlapping chunks from the same document. |
| Over budget | Relevant, but everything that fit the token budget ranked higher. | A general overview when a specific passage was also retrieved. |
| Cleaned | Noise was removed from inside the item; the item itself was kept. | 336 passing tests removed from a test run, both failures kept. |
Match the risk of the workload.
Not every agent should be optimized equally hard. The cost of a wrong answer differs by workload, so aggressiveness is set per agent, and the learned threshold adjusts around it.
| Workload | Setting | Why |
|---|---|---|
| Legal, compliance, medical | Conservative | A missing clause costs more than any token saving. |
| Customer support | Balanced | Answers depend on a few passages; the rest of top-k is usually noise. |
| Internal tooling | Aggressive | Users can retry, and volume makes every token count. |
Cut tokens, then verify the answer still had what it needed.
A decision model reads the question and the context that was kept, and answers one question: was it enough to answer? The response never waits for it.
Too thin, and that agent's relevance threshold drops so its next requests keep more. Comfortably enough, and it cuts a little more.
Every removed item is listed with its reason: duplicate, low relevance, over budget, or noise cleaned.
Lines that report an error, a failure or an exception survive every cut, whatever their relevance score.
Conservative, balanced or aggressive per agent. The learned threshold moves around that setting, within fixed bounds.
With the agent SDK, the full tool output stays in your process and the model can expand any omitted part.
Questions about quality guard
What does the quality check measure?
Whether the context that was kept was enough to answer the question completely and accurately, as a probability judged by a decision model after the response. It runs when something was removed for low relevance.
Does it slow down my requests?
No. It runs after the response is sent, together with the request log. Your request's latency does not include it.
Should I still keep an evaluation set?
Yes. The per-request check catches thin contexts as they happen; an evaluation set of real questions with known answers tells you whether answers stay right on your workload. They answer different questions.
Can removing context improve quality?
Often. Models use evidence buried in the middle of a long context less reliably than evidence near its edges (Lost in the Middle), so less noise can mean better answers, not just cheaper ones.
What happens if the decision model is unavailable?
Spendwaise never cuts your context on a guess. The API answers with a clear 503, and your code sends the original context as is: the request goes through, it just isn't optimized until the model is back.
Cut your LLM costs
without cutting quality.
One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.