The context optimization API for production LLM apps.
Keep your retrieval, keep your model. Add one call that returns the smaller, higher-signal context in milliseconds and tells you what it saved.
// Before your LLM call: send the context you would have sent
const res = await fetch(`${SPENDWAISE_URL}/api/v1/optimize`, {
method: "POST",
headers: { Authorization: `Bearer ${process.env.SPENDWAISE_KEY}` },
body: JSON.stringify({
query,
candidates: [...docs, ...toolOutputs], // { text, source }
agent: "support-triage", // per-agent budgets and stats
model: "claude-sonnet-5-5", // prices the savings
}),
});
const { context, report } = await res.json();
const answer = await llm.generate({ query, context });
console.log(report.tokens.savedPct, report.latencyMs); // 71.4 180Between retrieval and inference. Nowhere else.
The API is one call placed after you gather context and before you call your model. Your retriever, prompts and model stay exactly as they are. Spendwaise never calls your LLM and never sits in its billing path, so there is no model markup and no new point of failure for generation.
Rolling back is removing the call: pass your original context to the model again and the stack works as it did before.
A few inputs, two outputs.
The common case is a query and the context items you would have sent anyway. The response is the context to send and a report of what changed. Kept items come back in the order you sent them, never re-ranked, so a conversation reads the same on every turn.
| Field | Direction | Meaning |
|---|---|---|
query | Request | The current user question or task. Every context item is scored against it. |
candidates | Request | The context items: plain strings or { id, text, source }. RAG chunks, web pages, tool output, document sections. |
agent | Request, optional | Names the agent so budgets, aggressiveness, the learned threshold and stats apply per agent. |
budget | Request, optional | Maximum input tokens of context. See token budgets. |
model | Request, optional | The model you send to, so savings are priced at its input rate. |
conversation | Request, optional | Remembers every decision in a conversation, so an item sent again gets the same decision. See prompt caching. |
context | Response | The kept items, in the order you sent them, ready to pass to your model. |
report | Response | Tokens and cost before and after, per-item decisions with reasons, the threshold applied, latency and timings per step. |
The same call fits every context source.
Because every source becomes a list of context items, the integration looks the same whatever produced them. One rule holds for all of them: optimize new content once, when it first enters the conversation, and store what you sent.
Retrieve a wider top-k than you do today, optimize, then generate. Wider retrieval improves recall; the optimizer keeps the cost down.
Use the agent SDK, or optimize each new tool result before you append it to the messages.
Pass raw tool responses as context items: test runs, logs, JSON and HTML lose their noise before relevance is judged.
Optimize what you retrieve from long-term memory. Do not pass the conversation's own earlier turns back through: it would change history and break the prompt cache.
Pass fetched pages or passages and send the model only the ones that answer the question.
One call. Your stack stays the same.
Send a query and the context items you would have sent. Receive the context to send, in your order, and a per-item decision log.
Plain HTTPS and JSON. For agents, the agent SDK shrinks tool results inside a Vercel AI SDK agent.
Vector search results, web pages, MCP tool output and document sections all pass as context items.
The API never calls your LLM. It returns context you pass to OpenAI, Anthropic, Gemini or a local model.
Relevance by a small decision model, about 100 ms for most queries, every item scored in parallel. Each response reports its own latency, step by step.
Tokens and cost before and after, the threshold applied, timings and the removal reasons, ready to log.
Questions about api
Does Spendwaise call my LLM?
No. The API returns optimized context. You pass it to OpenAI, Anthropic, Gemini, an open model or your own model exactly as you do today.
Which languages are supported?
Any language that speaks HTTPS and JSON. For TypeScript agents on the Vercel AI SDK, the agent SDK (contextwaise on npm) wraps your model and tools in one line.
What counts as a context item?
Anything you would put in the prompt as context: a retrieved chunk, a web page or passage, a tool response or a document section.
How much latency does it add?
Relevance is judged by a small decision model, about 100 ms for most queries, with all items scored in parallel. API keys and project settings are cached, and the quality check and the request log run after the response. Every response reports latencyMs and per-step timings.
Will it break my prompt cache?
Not if you optimize new content once and store what you sent. Kept items return in your order, and a conversation id makes repeated items get the same decision. See prompt caching.
How do I remove it?
Delete the call and pass your original context to the model. Nothing else in your stack depends on Spendwaise.
Cut your LLM costs
without cutting quality.
One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.