Context engineering: a practical guide for production AI
Context engineering is the practice of deciding, for every model call, which information enters the context window and which does not. It replaces the habit of sending everything you have with a deliberate selection of the smallest high-signal token set. This guide defines the term, explains why it emerged, and gives a workflow you can apply to a RAG pipeline, an agent, or a chat product this week.
- Context engineering decides what reaches the model on each call. Prompt engineering decides how the instructions are phrased. The first is a systems discipline, the second is a writing discipline.
- Context is a finite resource with diminishing returns per token. Larger windows made it cheaper to be careless, not free.
- The core rule is the smallest high-signal token set: retrieve broadly, then send narrowly.
- Most waste comes from four places: unbounded chat history, raw tool payloads, over-fetched retrieval, and one giant system prompt.
- Measure tokens per request by source, cost, latency and quality on an evaluation set. Anything you cannot measure will grow back.
What is context engineering?
Context engineering is the discipline of assembling the right set of tokens for a language model call: instructions, tools, retrieved knowledge, memory, history and the current task, selected and sized so that the model has what it needs to answer and nothing that distracts it. The unit of work is the context window as a whole, not a single prompt string.
The term became common in 2025 as teams building agents noticed that prompt quality had stopped being the bottleneck. Anthropic's engineering guidance frames the goal as finding the smallest set of high-signal tokens that maximizes the likelihood of the desired outcome (Effective context engineering for AI agents). Vercel describes it as the practice of designing what information a model sees, and when (What is context engineering?). Microsoft's multi-agent reference architecture treats it as an architectural concern on par with orchestration and memory (Context engineering).
A useful mental model: prompt engineering writes the instructions; context engineering runs the supply chain that delivers everything else into the window, on a budget, on every call.
Why context engineering emerged
Context engineering emerged because context windows grew far faster than the model's attention budget and your inference budget did. A window that accepts 200,000 tokens does not attend to all of them equally, and it bills for all of them on every call.
Three shifts made this unavoidable. First, retrieval became easy: vector search, web search and MCP tool servers can return dozens of results in milliseconds, so the cheap default is to send all of them. Second, agents run in loops: each step appends tool output and reasoning to the transcript, so the context of step ten contains the debris of steps one through nine. Third, conversations became persistent: chat products keep history for weeks, and naive implementations replay all of it on every turn.
The result is a context window that is mostly filler. Models handle filler imperfectly. Relevant facts buried in the middle of a long context are recalled less reliably than facts near the start or end, and unrelated content increases the chance that the model anchors on the wrong detail. Quality degrades gradually rather than failing loudly, which is why teams often discover the problem through the invoice before they see it in the answers.
Context is therefore a finite resource with diminishing returns per token. The first 2,000 tokens of well-chosen evidence do most of the work. The next 20,000 add cost, latency and noise. Context engineering is the set of practices that keeps you on the right side of that curve. The cost side of this argument is covered in depth in the LLM cost optimization guide.
Anatomy of a context window
A production context window has seven or eight distinct parts, each with its own source, its own failure mode and its own budget. Treating them separately is the first step toward controlling them.
| Component | Where it comes from | Typical failure mode |
|---|---|---|
| System prompt | Static text written by your team | Grows into a wiki nobody prunes |
| Task instructions | Per-feature or per-route prompt | Duplicates the system prompt |
| Tool definitions | Function schemas registered with the model | Fifty tools loaded when three are relevant |
| Retrieved knowledge | Vector search, keyword search, web search | Top-k set too high, weakly related chunks included |
| Memory | User profile, preferences, past decisions | Everything ever stored is replayed |
| Conversation history | Prior turns in this session | Unbounded, includes resolved tangents |
| Tool results | API responses, database rows, file contents | Raw JSON with 90% irrelevant fields |
| Current user turn | The actual request | Often the smallest and most important part |
The components differ in volatility. The system prompt changes on deploy. Tool definitions change per feature. Retrieved knowledge, memory, history and tool results change on every call, and that is where context engineering spends most of its effort. The runtime, automated part of that work, deciding per call which retrieved items, memories, turns and tool results are worth sending, is what we call context optimization.
Core principles of context engineering
Six principles cover most of what works in practice. They are ordered from the one you apply most often to the one you apply only when the others are in place.
1. The smallest high-signal token set
Target the minimum context that lets the model answer correctly, not the maximum context that might help. This is a selection problem, so treat it as one: score each item against the current task, keep the ones that earn their tokens, and drop the rest. Selection is different from truncation. Truncation removes whatever is at the end. Selection removes whatever is least useful.
2. Just-in-time retrieval over pre-loading
Fetch information when a step needs it rather than loading it at the start in case a step needs it. An agent that reads a file when it is about to edit it uses a fraction of the context of an agent that reads the whole repository up front. Just-in-time retrieval keeps each call's window focused on the current step.
3. Progressive disclosure
Give the model a table of contents before giving it the book. Return document titles and summaries first, then sections, then passages, expanding only the branch the task follows. This is the same idea as lazy loading in software: pay for detail only where it is used. The RAG context compression guide walks through page to section to chunk narrowing.
4. Compaction
When a transcript or a tool result must stay in context but is too long, replace it with a faithful shorter version. Summarize resolved conversation turns. Reduce a 4,000-token API response to the twelve fields the task reads. Compaction is lossy by design, so it belongs behind a quality check rather than applied blindly.
5. Structured notes instead of raw memory
Keep memory as structured, addressable facts ("customer plan: Team", "prefers metric units") rather than as a pile of past transcripts. Structured notes can be retrieved selectively and cost tens of tokens each. Raw transcripts cost thousands and are retrieved wholesale. Redis's writing on agent context is a good reference on separating short-term working context from long-term stores (AI agent context).
6. Sub-agent context isolation
Give a sub-task its own clean context window and return only its result to the parent. A research sub-agent can read 80,000 tokens of sources and return a 500-token brief. The parent never carries the 80,000. Isolation is the most powerful lever for long-running agents and is discussed further in the agent context optimization guide.
Context engineering vs prompt engineering
Prompt engineering optimizes the wording of instructions inside a context window. Context engineering optimizes the composition of the entire window, including everything that is not an instruction. You need both, but they are different jobs with different tools.
| Dimension | Prompt engineering | Context engineering |
|---|---|---|
| Unit of work | A prompt string | The whole context window, per call |
| Main question | How should I phrase this? | What should be in here at all? |
| Changes when | You edit the prompt | Every request, based on task and budget |
| Owned by | Whoever writes prompts | The team that owns retrieval, memory and tools |
| Primary lever | Instructions, examples, format | Selection, ranking, budgets, compaction |
| Measured by | Task accuracy on fixed inputs | Tokens per request by source, cost, latency, quality on an eval set |
| Typical tooling | Prompt registry, A/B tests | Retrieval pipeline, ranker, budget enforcer, eval harness |
A practical test: if a change is made by editing text in a file, it is prompt engineering. If a change is made by editing code that decides what data is fetched, kept, ordered or dropped, it is context engineering.
A practical context engineering workflow
The workflow below takes an existing application from "we send whatever we retrieved" to "we send a budgeted, ranked, measured context" in six steps. Each step produces something you can look at before moving to the next.
- Inventory the sources. List every component that contributes tokens to a call: system prompt, tool schemas, each retrieval path, memory stores, history, and each tool that returns results. Most teams find one or two sources they had forgotten about.
- Measure tokens per source. Log the token count of each component on real traffic for a day. Report the median and the 95th percentile per source. This single table usually explains most of the bill and most of the latency.
- Set a budget per source and per call. Decide the ceiling before deciding the content. A support agent might get 1,500 tokens of history, 3,000 of retrieved evidence, 1,000 of tool results, and a total cap of 8,000. Budgets turn an open-ended problem into a packing problem. See token budgets for how to enforce them per agent.
- Select, deduplicate and rank. Score every item against the current task with something cheaper than the generation model: embedding similarity, a small classifier, or a cross-encoder. Collapse near-duplicates. Order survivors by expected contribution so the strongest evidence appears first.
- Compact what must stay. Summarize resolved history. Project tool results down to the fields the task reads. Trim retrieved passages to the sentences that carry the answer.
- Evaluate before and after. Run a fixed evaluation set through both pipelines. Compare answer quality and groundedness alongside tokens and cost. Ship only if quality holds. A quality guard turns this from a one-time check into a per-request constraint.
items = retrieve(query) # RAG, web, MCP, memory, history
scored = score(items, query) # cheap relevance decisions
kept = dedupe(scored) # collapse near-duplicates
ranked = rank(kept) # strongest evidence first
context = fit(ranked, budget=6_000) # whole items in, whole items out
guard(context, query) # refuse to over-cut
answer = llm.generate(query, context) # the only expensive callSteps one and two are a week of instrumentation. Steps three to six are the runtime layer, which you can build in-house or adopt. Spendwaise implements steps four to six as a single call between retrieval and inference and reports what it removed; the API page shows the shape.
A worked example with the arithmetic
Consider a customer-support agent that answers billing questions. The figures below are illustrative but the shape is typical of what the inventory step in the workflow reveals.
| Source | Before (tokens) | After (tokens) | What changed |
|---|---|---|---|
| System prompt | 1,800 | 1,100 | Removed duplicated policy text now retrieved on demand |
| Tool definitions | 2,400 | 600 | Load 4 billing tools instead of 22 |
| Retrieved knowledge | 6,200 (top-k 12) | 1,900 (4 kept of 20 retrieved) | Wider retrieval, stricter selection |
| Conversation history | 4,100 | 900 | Resolved turns summarized |
| Tool results | 3,600 | 500 | Invoice JSON projected to 9 fields |
| User turn | 320 | 320 | Unchanged |
| Total | 18,420 | 5,320 |
The reduction is 18,420 minus 5,320, which is 13,100 tokens, or 71% of input context per request. At an illustrative input price of $3 per million tokens, that is $0.0553 before and $0.0160 after, a saving of $0.0393 per request. At 200,000 requests a month the input-side saving is 200,000 times $0.0393, which is $7,860 per month. Latency falls with it because time to first token scales with prompt length.
Note what did not happen. Retrieval got wider, not narrower: 20 chunks fetched instead of 12, because the cost of fetching is negligible and the cost of sending is not. The user turn was untouched. Nothing was truncated mid-sentence. The saving came from selection, deduplication and compaction, each of which can be inspected. A savings report that shows this table per agent is how you keep the gains after the initial cleanup.
Common anti-patterns
Four patterns account for most of the wasted context in production systems. Each is a reasonable first implementation that was never revisited.
Dumping everything retrieved
Top-k retrieval returns k results whether or not results 5 through k are relevant. Sending all of them treats the retriever's ranking as a relevance decision, which it is not. Retrieval finds results. A second, cheaper judgement should decide which results earn a place in the window.
Unbounded conversation history
Replaying every prior turn is simple and wrong. Most turns are resolved, off-topic, or superseded. Keep the last few turns verbatim, summarize the rest, and retrieve older turns only when the current task refers to them.
Raw tool payloads
APIs return records designed for programs, not for prompts. A CRM contact record might carry 80 fields when the task reads three. Project tool results to the fields the task uses before they enter context, and summarize lists longer than the model needs to see.
One giant system prompt
System prompts accumulate rules, examples and edge cases until they cost thousands of tokens on every call, most of which do not apply to the current request. Move conditional guidance into retrievable documents and load it only for the routes that need it.
A fifth, quieter anti-pattern is optimizing cost without measuring quality. Aggressive trimming that drops the one passage the answer needed is worse than no trimming. Every reduction is a trade-off, so the goal is to reduce token spend without blindly reducing context.
How to measure context engineering
Context engineering is measured with four numbers per request: input tokens by source, cost, latency, and quality on a fixed evaluation set. Track them together, because any one of them can be improved by making the others worse.
| Metric | How to compute it | What a good trend looks like |
|---|---|---|
| Input tokens by source | Token-count each context component before the model call; log median and p95 | Falls after each change, with the largest sources shrinking most |
| Cost per request | Input tokens times model input price, plus output | Falls in proportion to input tokens |
| Latency | Time to first token and total time, p50 and p95 | Falls with prompt length; watch that selection overhead stays small |
| Answer quality | Score a fixed eval set with a rubric or judge, before and after | Flat or up; a drop means the selection is over-cutting |
| Groundedness | Fraction of claims supported by the sent context | Usually rises as noise leaves the window |
| Guard interventions | Count of requests where the quality check kept extra context | Low and stable; a spike flags a retrieval or budget problem |
Report these per agent, per source and per customer. Aggregate numbers hide the one agent that dumps its whole tool history into every call. The evaluation set should be small, fixed and representative: 50 to 200 real queries with reference answers is enough to detect a regression, and it must run on every change to selection logic, not only on the first one.
Where context optimization fits
Context optimization is the runtime, automated part of context engineering: the layer that scores, deduplicates, ranks, budgets and guards context on every call, after retrieval and before inference. The rest of context engineering is design-time work on prompts, tools, memory schemas and agent structure.
The split matters because the two halves change at different speeds. Design-time decisions change when you ship a feature. Runtime decisions change on every request, because the right context for "why was I charged twice" differs from the right context for "cancel my plan" even inside the same agent with the same prompt. Automating the runtime half is what makes the design-time half stay true under real traffic.
You can build the runtime layer yourself from an embedding model, a small classifier and a packing routine, and many teams should start there. Adopt a dedicated layer such as Spendwaise when you need the selection to be model-agnostic across RAG, web, tools and memory, when you want per-agent budgets and a quality guard enforced centrally, and when finance needs a report of what was removed and what it saved. Either way, the principles in this guide are the same.
Frequently asked questions
What is context engineering?
Context engineering is the practice of deciding which information enters a language model's context window on each call, and in what form. It covers instructions, tool definitions, retrieved knowledge, memory, conversation history and tool results, and it aims for the smallest set of high-signal tokens that lets the model answer correctly.
What is the difference between context engineering and prompt engineering?
Prompt engineering optimizes the wording of instructions. Context engineering optimizes the composition of the whole context window, including everything that is not an instruction: what is retrieved, what is remembered, what history is replayed and what tool output is kept. Prompt engineering is done by editing text; context engineering is done by changing what data is fetched, selected, ordered and dropped.
Is context engineering the same as RAG?
No. RAG is one source of context. Context engineering governs all sources, including tools, memory and history, and it governs what happens to retrieved chunks after retrieval: selection, deduplication, ranking and budgeting. A RAG pipeline with no selection step is a common example of missing context engineering.
Why not just use a bigger context window?
A bigger window raises the ceiling but does not lower the cost or improve attention. Every token is billed and every token competes for the model's attention. Relevant facts buried among irrelevant ones are recalled less reliably, so a larger window filled carelessly often produces worse answers at higher cost than a smaller window filled deliberately.
How do I measure whether context engineering is working?
Track input tokens per request broken down by source, cost per request, latency, and answer quality on a fixed evaluation set, before and after each change. Tokens, cost and latency should fall while quality holds flat or rises. If quality falls, the selection is over-cutting and the budget or the relevance threshold needs adjusting.
What is context optimization?
Context optimization is the runtime, automated part of context engineering: a layer that sits between retrieval and inference and, on every call, scores context for relevance, removes duplicates, ranks what remains, fits it to a token budget and checks that the cut did not remove what the answer needs. It is what turns context engineering principles into per-request behavior.