Spendwaise
Quality guard

Context compression with a quality check on every request.

Every reduction is a trade-off. Spendwaise checks each one after the response, learns from it, and shows you every decision it made.

See it live
What was removedrequest 8f3a…
duplicate of #3-412 tok
Refund requests are processed within 5 to 7 business days…
low relevance-288 tok
Our office hours are Monday to Friday, 9am to 6pm CET…
cleaned-7,053 tok
336 passing tests and 6 progress bars removed from the test run…
over budget-630 tok
The 2023 pricing page listed three plans: Starter, Team…
Why a guard

A cost report can look perfect while answers get worse.

Every context reduction technique can be tuned to show a bigger saving. Cut harder and the token count falls further. The problem is that the cost dashboard improves in exactly the same way whether the removed context was noise or the one paragraph the answer depended on. Nobody notices the difference until users do.

Aggressive compression fails in three recognizable ways: it over-cuts the evidence an answer needs, it loses the passages the model would have cited, and it breaks groundedness by feeding the model paraphrases instead of sources. The RAG context compression guide covers each failure in detail. The quality check exists so that none of them can happen silently.

How it works

A check after every response, a correction before the next.

The check runs only when the relevance threshold actually removed something, and only after your response has been sent. It turns every request into feedback for the agent's next one.

01
Select

The request is optimized at the agent's current threshold and the context is returned to you.

02
Respond

Your model gets the context. Nothing below adds a millisecond to that response.

03
Check

A decision model answers one yes/no question about the kept context: is it enough to answer the question completely and accurately? The answer is a calibrated probability.

04
Adjust

Below 50%, the agent's threshold drops by 0.15 so it keeps more. At 85% or above, it rises by 0.05 so it cuts more. In between, it stays.

05
Bound

The threshold stays within fixed limits, and steps down are three times larger than steps up: quality recovers faster than savings grow.

The removal log

Every removed item comes with a reason.

Trust in an optimizer comes from being able to audit it. Each request records what was removed and why, so a surprising answer can be traced back to the context decision behind it.

ReasonWhat it meansTypical example
Low relevanceThe item scored below the agent's threshold for this question.A billing policy retrieved for a shipping question.
DuplicateThe same information is already in the context.Two overlapping chunks from the same document.
Over budgetRelevant, but everything that fit the token budget ranked higher.A general overview when a specific passage was also retrieved.
CleanedNoise was removed from inside the item; the item itself was kept.336 passing tests removed from a test run, both failures kept.
Aggressiveness

Match the risk of the workload.

Not every agent should be optimized equally hard. The cost of a wrong answer differs by workload, so aggressiveness is set per agent, and the learned threshold adjusts around it.

WorkloadSettingWhy
Legal, compliance, medicalConservativeA missing clause costs more than any token saving.
Customer supportBalancedAnswers depend on a few passages; the rest of top-k is usually noise.
Internal toolingAggressiveUsers can retry, and volume makes every token count.

Cut tokens, then verify the answer still had what it needed.

Checked after every response

A decision model reads the question and the context that was kept, and answers one question: was it enough to answer? The response never waits for it.

Self-tuning threshold

Too thin, and that agent's relevance threshold drops so its next requests keep more. Comfortably enough, and it cuts a little more.

Removal log

Every removed item is listed with its reason: duplicate, low relevance, over budget, or noise cleaned.

Errors always kept

Lines that report an error, a failure or an exception survive every cut, whatever their relevance score.

Tunable aggressiveness

Conservative, balanced or aggressive per agent. The learned threshold moves around that setting, within fixed bounds.

Nothing deleted

With the agent SDK, the full tool output stays in your process and the model can expand any omitted part.

Questions about quality guard

What does the quality check measure?

Whether the context that was kept was enough to answer the question completely and accurately, as a probability judged by a decision model after the response. It runs when something was removed for low relevance.

Does it slow down my requests?

No. It runs after the response is sent, together with the request log. Your request's latency does not include it.

Should I still keep an evaluation set?

Yes. The per-request check catches thin contexts as they happen; an evaluation set of real questions with known answers tells you whether answers stay right on your workload. They answer different questions.

Can removing context improve quality?

Often. Models use evidence buried in the middle of a long context less reliably than evidence near its edges (Lost in the Middle), so less noise can mean better answers, not just cheaper ones.

What happens if the decision model is unavailable?

Spendwaise never cuts your context on a guess. The API answers with a clear 503, and your code sends the original context as is: the request goes through, it just isn't optimized until the model is back.

Cut your LLM costs
without cutting quality.

One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.

See it live
Keep your model, retriever and prompts. Remove one call to roll back.