Agent context optimization: how to stop context bloat in multi-step agents
Agent context optimization is the practice of deciding, at every step of an agent loop, which tool outputs, memories and retrieved documents are worth sending to the model. Without it, context compounds across steps and most of the tokens the model reads never contribute to the task. The fix is to filter, summarize, deduplicate and budget context before each model call, then verify that task success holds.
- Agent context grows with every step because each model call re-reads the full history, including every previous tool result. A 10-step run can read the same payload ten times.
- Most agent waste comes from four places: raw JSON tool payloads, MCP responses, stale intermediate state and repeated retrievals of the same content.
- Six patterns fix it: filter tool output before it enters context, summarize completed sub-tasks, retrieve memory by task instead of replaying history, set a per-step token budget, deduplicate across steps and keep only task-relevant state.
- Optimize context before inference, not after. Cheap relevance decisions on context cost far less than paying the generation model to read it.
- Measure both sides: input tokens per step and task success rate on a fixed evaluation set. A reduction that lowers success is not a saving.
Why agent context grows with every step
Agent context grows because every model call in an agent loop re-reads the system prompt, the full conversation so far and every tool result that has been appended since the run started. Nothing is removed by default, so the prompt for step N contains the output of steps 1 through N minus 1.
A single chat request has a fixed prompt. An agent does not. Each iteration adds a model message, one or more tool calls and their results, and often a fresh retrieval. Because most agent frameworks keep the message list as the state, the entire list is serialized into the next request. The cost is quadratic in the number of steps: the total tokens read across a run is the sum of an increasing sequence, not a flat rate per step.
The arithmetic for a 10-step run
Take an illustrative agent with a 2,000-token system prompt and tools that return about 3,000 tokens per call. If nothing is filtered, the context at each step is the system prompt plus everything accumulated so far.
| Step | Context read by the model | Cumulative tokens read |
|---|---|---|
| 1 | 2,000 + 0 = 2,000 | 2,000 |
| 2 | 2,000 + 3,000 = 5,000 | 7,000 |
| 3 | 2,000 + 6,000 = 8,000 | 15,000 |
| 5 | 2,000 + 12,000 = 14,000 | 40,000 |
| 10 | 2,000 + 27,000 = 29,000 | 155,000 |
Ten steps that produced 27,000 tokens of tool output cause the model to read 155,000 input tokens in total. The first tool result is read nine times. If the same run kept each step to a 6,000-token budget, the total would be at most 10 × 6,000 = 60,000 tokens, and the model would still see the parts of each result that mattered. These numbers are illustrative; the shape of the curve is not.
Where agent context waste comes from
Agent context waste comes from information that enters the prompt in the shape it arrived in, rather than in the shape the task needs. Five sources account for most of it.
Raw JSON tool payloads
A CRM lookup returns 40 fields per record when the agent needs three. A search API returns ranking metadata, snippets, URLs and timestamps for 20 results when the agent needs two. Serialized JSON is also token-expensive: keys, quotes, braces and whitespace routinely double the token count of the useful values.
MCP responses
Model Context Protocol servers expose tools that were designed for completeness, not for prompt size. A file listing, a database query or an issue tracker search often returns hundreds of rows because the server cannot know what the agent will do with them. The agent framework appends the whole response to context.
Stale intermediate state
Once a sub-task is finished, the reasoning that produced it is no longer useful. The plan the agent drafted at step 1, the three failed attempts at step 4 and the raw data behind a number it has already computed all stay in the message list and are re-read on every later step.
Repeated retrievals
Multi-step agents retrieve more than once. Each retrieval for a related query returns overlapping chunks, and the overlap is appended again. The same product page, policy paragraph or log excerpt ends up in context three or four times with slightly different boundaries.
Long chat history
In assistants that run over a long session, turns from an hour ago are still present in every request. Most of them concern tasks that are complete. Anthropic's guidance on effective context engineering treats context as a finite resource with diminishing returns, and history is the source that most often exceeds its useful size.
Source, typical waste and the fix
Each context source in an agent has a characteristic form of waste and a matching fix. The table below is the map most teams end up drawing after their first cost review.
| Context source | Typical waste | Fix |
|---|---|---|
| Tool output (APIs, SQL, CRM) | Full records with dozens of unused fields, verbose JSON encoding | Project to the fields the task needs; render as compact text before appending |
| MCP responses | Hundreds of rows or files returned for completeness | Filter or summarize records against the current task before they enter context |
| Intermediate state | Plans, drafts and failed attempts from finished sub-tasks | Replace completed sub-task transcripts with a short result summary |
| Retrieved documents | Overlapping chunks from related queries across steps | Deduplicate across the run; keep one representative per fact |
| Conversation memory | Old turns replayed in every request | Retrieve only the history relevant to the current step |
| Web search | Whole pages when a paragraph answers the question | Select relevant passages, not pages |
Every row is the same move: decide what is worth sending before inference, using something much cheaper than the generation model. The broader case for that layer is covered in LLM context optimization; the retrieval-specific version is in RAG context compression.
Six patterns that fix agent context bloat
Agent context bloat is fixed by inserting a decision step between the sources that produce context and the model call that consumes it. Six patterns cover almost every agent, and they compose.
1. Filter tool output before it enters context
Treat every tool result as optional context, not as a fact the model must see. Score each record, row or field against the current step's goal and keep the subset that can change the model's next action. A search result with a relevance score below the threshold is dropped before serialization. This is the single highest-value pattern because tool output is both the largest source and the easiest to judge.
2. Summarize completed sub-tasks
When a sub-task closes, replace its transcript with a compact result: what was asked, what was found, what was decided. The reasoning and raw data are stored outside the prompt and can be re-fetched if a later step needs them. A 4,000-token investigation becomes a 120-token note.
3. Retrieve memory by task, not by recency
Do not replay the full session. Store turns and facts in a memory store and retrieve the ones relevant to the current step, the same way you retrieve documents. Redis describes this separation of short-term working context from long-term memory in its overview of AI agent context. The prompt then contains the last few turns plus what the current task needs, instead of everything that ever happened.
4. Set a per-step token budget
Give each step a ceiling for input context and fill it by value per token, not by arrival order. A budget turns a variable, compounding cost into a predictable one and forces the selection step to exist. Whole items should be kept or dropped; truncating the tail of a message is the wrong tool. Token budgets explains how to size them per agent.
5. Deduplicate across steps
Before appending a new retrieval or tool result, compare it with what is already in context. Near-identical chunks collapse into one. Records already summarized are not re-added. Dedup is cheap with embeddings or hashing and is the main defense against repeated retrievals.
6. Keep only task-relevant state
Build the prompt for each step from a state object, not from an append-only log. The state holds the goal, the current plan, the open sub-tasks and the evidence still needed. Serialization happens from the state, so anything that is no longer relevant simply stops being included. This is the shift from message history as state to state as the source of the prompt.
An agent loop with a context optimization step
The change to an agent loop is one step: after collecting context and before calling the model, select what fits the budget and the task. Everything else stays the same.
state = { goal, plan: [], evidence: [], done: [] }
for step in range(MAX_STEPS):
# 1. gather context from every source
items = []
items += memory.retrieve(query=state.goal, k=20)
items += state.evidence
items += last_tool_results # raw payloads, not yet filtered
# 2. optimize context before inference
context, report = optimize(
query=current_subtask(state),
items=items,
budget_tokens=6_000,
dedupe=True,
quality_guard=True, # keep items the answer depends on
)
log(report.tokens_before, report.tokens_after, report.dropped)
# 3. call the model with the optimized context
action = llm.generate(system=SYSTEM, state=state.summary(), context=context)
# 4. execute, then store results outside the prompt
last_tool_results = run_tools(action)
state.evidence += project_fields(last_tool_results) # keep useful fields only
if action.closes_subtask:
state.done.append(summarize(action)) # replace transcript with a note
if action.is_final:
return action.answerTwo details matter. First, the optimize call takes the current sub-task as the query, not the original goal, because relevance changes as the agent progresses. Second, tool results go into state.evidence after field projection, so the raw payload is never serialized into a later prompt. If you want a hosted version of the optimize step, the Spendwaise API returns the optimized context plus a report of what was removed and why.
What it does to latency and cost
Optimizing agent context lowers both cost and latency, because the model reads fewer input tokens on every step and the selection step is cheaper than the tokens it removes.
Cost follows input tokens directly. Using the illustrative 10-step run above, moving from 155,000 to at most 60,000 tokens read is a 61% reduction in input tokens for that run. At a hypothetical input price of $3 per million tokens, that is $0.465 versus $0.180 per run. Multiply by runs per day to see why agents dominate LLM bills faster than chat does. The full cost picture, including output tokens, caching and model choice, is in the LLM cost optimization guide.
Latency follows input tokens too. Time to first token grows with prompt length, and in a 10-step run that delay is paid ten times. Relevance decisions on context run in parallel on small models and typically complete well before the generation model would have finished reading the tokens they removed. The net effect is a faster run, not a slower one, provided the optimization step is genuinely cheap.
- Fewer input tokens per step: the direct saving, compounding across steps.
- Lower time to first token per step: the model starts generating sooner.
- Fewer context-window overflows: long runs stop failing at step 12 because the history no longer fits.
- Better answers on late steps: less noise in the prompt means the model attends to the evidence that matters.
How to evaluate without breaking task success
Evaluate agent context optimization by measuring task success on a fixed set of runs before and after, alongside input tokens per step. Cost is the objective and quality is the constraint, never the other way around.
- Build an evaluation set of 50 to 200 real tasks with a known correct outcome or a rubric. Include long runs; short runs hide the compounding.
- Record the baseline: success rate, steps per run, input tokens per step, total tokens per run, wall-clock time.
- Turn on optimization with a conservative budget and a quality guard that keeps any item the current sub-task explicitly references.
- Re-run the set. Compare success rate first. If it dropped, inspect the removal log for the failed runs and find which drop caused it.
- Tighten the budget in steps only while success holds. Stop at the point where the next step costs accuracy.
- Keep the evaluation running in production on a sample of traffic, so drift in tools or data shows up as a quality alert rather than a support ticket.
The removal log is the debugging tool. Every dropped item should carry a reason: duplicate, off-topic, stale, over budget. When a run fails, the question is whether a needed item was dropped, and the log answers it in seconds. The quality guard exists to make that failure rare, and the savings report puts tokens saved next to quality so nobody optimizes blind.
Where agent context optimization fits in your stack
Agent context optimization sits between your tools, memory and retrieval on one side and the model call on the other. It does not replace the agent framework, the vector database or the LLM gateway.
| Layer | What it does | Relationship to context optimization |
|---|---|---|
| Agent framework | Runs the loop, dispatches tools, holds state | Hosts the optimize step; provides the context and the current sub-task |
| Vector database and retriever | Finds matching documents | Feeds results; a wider top-k becomes safe because selection happens after |
| LLM gateway | Routes, caches, tracks spend | Sees fewer input tokens; caching and routing still apply |
| Observability | Traces calls, shows cost per step | Displays the before and after tokens and the removal reasons |
The wider discipline this belongs to is context engineering: treating the prompt as an engineered artifact assembled from sources, rather than a log that happens to be sent. Agents are where that discipline pays back fastest, because they are where context compounds. Send fewer, better tokens at every step and the rest of the stack gets cheaper and faster for free.
Frequently asked questions
What is agent context optimization?
Agent context optimization is the practice of selecting, at each step of an agent loop, which tool outputs, memories and retrieved documents are sent to the model. It filters, deduplicates, summarizes and budgets context before inference so the model reads fewer, more relevant tokens without losing what the task needs.
Why do agents use so many more tokens than chat?
Because each step re-reads the full history, including every previous tool result. Tokens are paid for on every subsequent step, so total input tokens grow roughly with the square of the number of steps. A 10-step run with 3,000-token tool results reads about 155,000 tokens if nothing is filtered.
Is truncating the message history enough?
No. Truncation drops the oldest content regardless of relevance, which often removes the goal or a key fact while keeping recent noise. Selecting by relevance to the current sub-task, with whole items kept or dropped, preserves what matters and removes what does not.
Does filtering tool output slow the agent down?
Usually the opposite. Relevance decisions run on small models in parallel and finish faster than the generation model would take to read the tokens they remove. Time to first token drops on every step, and the run finishes sooner.
How do I know optimization is not hurting answer quality?
Run a fixed evaluation set before and after, compare task success rate first and tokens second, and keep a removal log with a reason for every dropped item. Tighten the budget only while success holds. In production, sample traffic into the same evaluation so drift shows up as an alert.
Do I need a separate product to do this?
No. Field projection, summaries of completed sub-tasks, dedup and a per-step budget can all be built into your loop. A service such as Spendwaise packages the selection, budgeting, quality guard and reporting behind one call if you would rather not maintain it, but the patterns work either way.