Context optimization that never breaks your prompt cache.
Providers bill a repeated prompt start at a fraction of the price, but only while it is byte-identical. Spendwaise trims new context once and never edits what the model already saw.
A discount on the part that did not change.
LLM APIs are stateless: every turn resends the whole conversation. Providers cache what they have already processed and bill the repeated start of a prompt at a fraction of the input price: Anthropic charges about a tenth for cache reads and a quarter more for writes (Anthropic docs), OpenAI and Gemini apply their discounts automatically.
The catch is in one word: prefix. The cache is matched byte by byte from the start of the prompt. At the first byte that differs, everything after it is billed at full price again, and written to the cache again.
Most context compression edits the past.
Compression that runs on the whole conversation before every call rewrites what earlier turns contained: a chunk kept at turn 3 is dropped at turn 4 because the question changed, a tool result is summarized after the fact, old turns are cleared to make room. Each edit moves the first differing byte to the middle of the history, and the cache is lost from there on. On a long conversation that can cost more than the trimming saved.
On newer Claude models there is a second cost. The model's saved reasoning is tied to the exact conversation that produced it, and editing an earlier turn invalidates it; for accounts created since late August 2026 the request is rejected outright.
| Cache killer | What happens | Spendwaise |
|---|---|---|
| Re-trimming earlier turns | History changes, cache lost from that point | Never: decisions are made once and remembered |
| Re-ranking context | Items move, the prefix changes | Kept items stay in your order |
| Switching models mid-conversation | Caches are per model, all lost | Routing decides once per conversation |
| Timestamps in the system prompt | The very start changes on every call | Not ours to fix, but our checklist flags it |
A filter at the door, not an editor of the past.
Spendwaise optimizes each piece of context the first time it enters the conversation: the chunks retrieved for this turn, the tool result that just came back. Your app stores the optimized version and resends it unchanged. From the provider's point of view, the history only ever grows at the end, which is exactly what its cache is built for.
If your code rebuilds the context from the original sources on every turn and you cannot change that, send a conversation id with the API: every decision in that conversation is remembered, and a resent item gets the same decision again, so the history the model sees still never changes.
A short checklist for a warm cache.
Spendwaise keeps its own output stable. These keep the rest of your prompt stable too.
No dates, user names or modes in it. Put changing information at the end of the conversation instead.
Tool definitions sit at the start of the prompt. Adding, removing or reordering one resets the whole cache.
Store what you sent and resend it unchanged. Pass only new content through Spendwaise.
Anthropic needs an explicit cache_control breakpoint; OpenAI and Gemini cache automatically.
Caches expire after minutes of inactivity. Long pauses between calls mean paying to write the cache again.
Cache-safe by design.
Each piece of context is optimized when it first enters the conversation. What was sent is never re-optimized.
A tool result's view is decided once and replayed byte for byte by tool call id. Cleanup is deterministic.
Kept items come back in the order you sent them, never re-ranked, so turn 3 reads the same on turn 30.
Pass a conversation id and every decision is remembered, so an item sent again gets the decision it got the first time.
Model routing decides once per conversation and only advises a switch when it pays back the lost cache.
Newer Claude models reject edited history with an error. Spendwaise never produces one.
Questions about prompt caching
Does Spendwaise ever change a message that was already sent?
No. It only decides what new context to send. With a conversation id it also remembers its decisions, so content sent again is decided the same way.
Does the agent SDK keep the cache?
Yes. A tool result is shrunk once, when the tool returns, and that view is stored in the conversation. If the framework asks for it again, the same view is replayed byte for byte.
Is caching a replacement for context optimization?
No. Caching discounts the repeated prefix; it does nothing for the new content each turn adds, and the first read is always at full price. Cut variable context first, then cache the stable prefix.
What about provider-run tools like web search?
Tools the provider runs itself never pass through your code, so nothing can trim them. Cap them with the provider's own settings, such as the maximum number of searches or fetched tokens.
Cut your LLM costs
without cutting quality.
One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.