Spendwaise
Token budgets

Reduce LLM input tokens with a budget, not a truncation.

A budget is a ceiling on what your model reads. Spendwaise fills it with the highest-signal context instead of cutting from the end.

See it live
Token budgetsper agent · input tokens
support-triage5,260 / 6,000
research-planner9,840 / 12,000
sql-analyst8,000 / 8,000
email-drafter3,120 / 4,000
doc-translator9,600 / 10,000
The problem

Truncation is a budget without judgment.

Most LLM applications already cap their input. They lower top-k, keep the last N chat turns, or cut the prompt at a character limit. Each of these is a token budget, and each fills it the same way: by position. Whatever arrived first survives, whatever arrived last is chopped, and nobody checks whether the chopped part was the one the answer needed.

A Spendwaise token budget keeps the ceiling and fixes the choice. The limit is still a hard number of input tokens, but the context inside it is selected by relevance to the current request, so a budget removes the least useful items first instead of the most recent ones.

How a budget is filled

Value per token, not order of arrival.

Every request passes through the same five steps. The budget is the last constraint applied, after the layer already knows what each item is worth.

01
Score

Each context item is scored against the current query: how much can it contribute to this answer?

02
Count

Each item's token cost is measured, so a long, weakly relevant document is compared fairly with a short, precise one.

03
Rank

Items are ordered by value per token. A 200-token passage that answers the question beats a 3,000-token page that mentions it.

04
Pack

Whole items are added in rank order until the budget is full. Nothing is sliced mid-passage.

05
Log

Every item that did not fit is recorded with its score, so you can see what the budget cost you.

Choosing a budget

Start from what you send today, then tighten.

The right budget depends on the task, not on the model's context window. A good first budget is the median input size the agent uses today. Lower it in steps and watch answer quality on your evaluation set; stop when quality moves. The quality guard automates that check.

The table below gives illustrative starting points. They are places to begin tuning, not recommendations for your workload.

WorkloadIllustrative starting budgetWhy
Support triage4,000 to 8,000 tokensAnswers usually depend on a few policy passages and the recent turns.
SQL and structured tool use6,000 to 10,000 tokensSchema and a handful of example rows matter; full result sets rarely do.
Research and synthesis10,000 to 20,000 tokensAnswers combine several sources, so breadth is worth paying for.
Agent steps after the firstSmaller than the first callLater steps need task state and the last tool result, not the whole history.
Budget and cost

A budget turns input cost into a number you can plan.

Input cost per request is input tokens times the model's input price. Without a budget, input tokens drift upward as retrieval, memory and tool output grow. With one, the input side of every request has a ceiling.

Illustrative: an agent that sends 18,000 input tokens per request at $3 per million input tokens spends $0.054 per request on input. A 6,000-token budget caps that at $0.018. At one million requests a month, that is the difference between $54,000 and $18,000 of input spend. The LLM cost optimization guide walks through the full cost equation.

Predictable input tokens, predictable cost.

Per-agent budgets

Give support-triage 6,000 tokens and research-planner 12,000. Each request is fitted to its own limit.

Budget-aware selection

Items are chosen by value per token, so a budget removes the least useful context first.

No silent truncation

Nothing is chopped mid-passage. Whole items are kept or dropped, and you can see which.

Cost forecasting

A budget caps input tokens, so per-request cost becomes a number you can plan around.

Latency control

Fewer input tokens means faster time to first token on every model.

Overrides when needed

Raise the budget for a request that needs it. The default stays tight.

Questions about token budgets

What is a token budget?

A token budget is a hard limit on how many input tokens of context a request may send to the model. Spendwaise fills that limit with the most useful context for the current request instead of cutting by position.

Is a token budget the same as max_tokens?

No. max_tokens limits how many tokens the model may generate in its answer. A token budget limits how much context goes into the request. They control opposite sides of the bill and are usually set together.

What happens when the context an answer needs does not fit?

The quality guard protects against over-cutting: if dropping context would starve the answer, it is kept and the request is flagged. You can also raise the budget for a single request that needs it while the default stays tight.

Can different agents have different budgets?

Yes. Budgets are set per agent, so a support agent and a research agent in the same product each get a limit that matches their task. Per-agent usage appears in the savings report.

Does a smaller budget always mean lower quality?

No. Removing irrelevant context often improves answers, because the model has less noise to read past. Quality only drops when the budget forces out evidence the answer depends on, which is what the evaluation set is there to catch.

Cut your LLM costs
without cutting quality.

One layer between retrieval and inference removes the context your model never needed, in milliseconds. Fewer input tokens, lower cost, a prompt cache that keeps hitting, and quality checked on every request.

See it live
Keep your model, retriever and prompts. Remove one call to roll back.