Shrink agent tool results the moment they return.
Agents replay every tool result on every later step. The Spendwaise SDK (contextwaise on npm) cuts each one once, before the model reads it, in one line for the Vercel AI SDK.
import { generateText, isStepCount } from "ai";
import { openai } from "@ai-sdk/openai";
import { createContextwaise } from "contextwaise";
import { withContextwaise } from "contextwaise/ai-sdk";
const cw = createContextwaise({ agent: "code-review" });
const { model, tools } = withContextwaise(cw, {
model: openai("gpt-6.1-sol"),
tools: { runTests, readFile, searchLogs },
});
// Every tool result reaches the model already shrunk.
// The full output stays with you, one expand() call away.
const { text } = await generateText({
model, tools, prompt, stopWhen: isStepCount(10),
});A tool result is paid for on every step after it.
In an agent loop, every model call resends the whole conversation, tool results included. A 20,000-token test log read at step 3 of a 20-step run is billed again on each of the 17 steps that follow. Prompt caching makes those repeats cheaper, not free, and the first read is always at full price.
Most agents need a small part of what their tools return: the two failing tests, not the 336 passing ones; the one folder that matches, not the whole tree. On our demo agent, a 7,209-token test run reached the model as 156 tokens, and a 28,309-token log as 564, with every error kept.
Cleaned, judged, then sent. Once.
Every tool result goes through the same steps the moment the tool returns. Most results never need the network: cleanup alone brings them under budget.
A tool call already seen gets exactly the same view as before, byte for byte.
Noise is removed by content type: test output, logs, JSON, HTML and file listings (grouped by folder, with nothing lost).
A result that fits its budget passes whole. Give file-reading tools a bigger budget so files the agent may edit arrive intact.
Bigger results are cut into parts and scored against the agent's intent by our decision model through the API, or locally without a key.
The relevant parts reach the model in their original order, with a note saying what was omitted and how to expand it.
What the model reads instead.
A Vercel AI SDK agent on OpenAI GPT-6.1 Sol with six synthetic tools, five questions in one conversation. Tools returned 52.6K tokens; the model read 9.1K. The agent's cost was $0.057 instead of about $0.129 (the comparison replays the same requests with raw outputs through a prompt-cache simulator). Your workload will differ; the dashboard measures yours.
| Tool output | Returned | Model read | What happened |
|---|---|---|---|
| Test run | 7,209 tokens | 156 tokens | 336 passing tests and progress bars removed, both failures kept. |
| Application logs | 28,309 tokens | 564 tokens | Repeated lines collapsed, then the relevance model kept the errors, their causes and warnings. |
| File listing | 3,782 tokens | 481 tokens | 453 paths grouped into 28 folder lines. No path lost. |
| Orders API (JSON) | 3,450 tokens | 1,408 tokens | 216 empty fields dropped. Every order kept: a total needs all of them. |
Vercel AI SDK today. Your stack next.
The SDK plugs into the AI SDK's own toModelOutput hook, the place the framework designed for giving the model a different view of a tool result than the app gets. The same core works with any framework that lets you see a tool result before the model does.
Adapters for LangChain, the OpenAI Agents SDK, the Claude Agent SDK and MCP are next. Until then, any agent can call the API on each new tool result before appending it to its messages.
Built for agent loops.
withContextwaise(cw, { model, tools }) wraps your Vercel AI SDK model and tools. Your app still gets every raw tool output.
Passing tests, progress bars, repeated log lines, empty JSON fields and HTML markup go, locally, in 1 to 20 ms on our demo agent.
Relevance is judged against what the agent is doing: the user's request and what the model said before calling the tool.
Omitted parts stay in your process. The model calls expand with a question and gets back exactly the part it needs.
Each result's view is decided once and replayed byte for byte by tool call id, so the provider's prompt cache keeps hitting.
Only sections that need a relevance decision go to the API. Dashboard reports carry sizes and tool names, never tool output.
Questions about agent sdk
Does the model lose information?
No. Omitted parts stay in your process, and every trimmed result tells the model how to get them back: it calls expand with a question and receives exactly the part it needs. Results that fit their budget, and anything reporting an error, are never cut.
Is it safe for prompt caching?
Yes. A tool result is shrunk once, when it returns, and that view is what your conversation stores. If the framework asks again, the same view is replayed byte for byte. See prompt caching.
Do I need an API key?
No. Without one, the SDK still cleans tool output locally and free, then sends it whole. With a Spendwaise key, relevance is judged by our decision model, only what the task needs is kept, and every result appears on your dashboard.
What leaves my machine?
Without a key, nothing. With one, only the parts of results that need a relevance decision, and usage reports with sizes and tool names. Full outputs never leave your process.
What if the API is slow or down?
The agent never waits on it: scoring times out after 4 seconds and for the next 30 seconds the model reads cleaned outputs whole, nothing cut. Reports are sent in the background and dropped if they cannot be.
Cut your LLM costs
without cutting quality.
One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.