--- title: "Context engineering for AI agents" description: "Context engineering selects the instructions, tools, evidence, history and state included in each model call. Learn how to manage and measure that context." canonical: https://past.dev/context-engineering last-updated: 2026-10-09 --- # Context engineering for AI agents Source: https://past.dev/context-engineering Context engineering selects the content for each model call: instructions, tool definitions, retrieved evidence, conversation history, working state and the output contract. Prompt engineering covers the wording of instructions. Context engineering also manages retrieval, compression, isolation, persistence and the token budget. Memory supplies information from previous calls and sessions. ## What is context engineering? Each model call has a token budget. Context engineering determines what is included, summarized or omitted, and the order of the included content. Agent applications assemble context in code at every step. The available sources and history can grow during a run, so the application must select content repeatedly. Anthropic's [Effective context engineering for AI agents](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) provides additional guidance. ## Context engineering vs prompt engineering Prompt engineering optimizes the wording of instructions. Context engineering determines every component of the context window and its token allocation. | | Prompt engineering | Context engineering | | --- | --- | --- | | Unit of work | The wording of an instruction | The composition of the whole window | | Who does it | A person, once, at design time | Code, on every call, at runtime | | Changes when | The task changes | The application state changes | | Typical failure | The model misunderstands the instruction | The required evidence was absent or received too little attention among irrelevant context | | Fixed by | Rewriting the instruction | Retrieving better, compressing harder, or splitting the task | | Measured by | Output quality on a fixed set of prompts | Tokens per call, cost per answer, and how often the needed evidence was present | ## What competes for the window Six types of content consume the token budget. Measure each type separately because they grow at different rates. - **System instructions.** Usually fixed for a given agent version. - **Tool definitions.** Fixed per call and proportional to the number and complexity of available tools. - **Retrieved evidence.** Variable and controlled by retrieval limits and ranking. - **Conversation history.** Grows during a session unless summarized or truncated. - **Tool results.** Can consume a large part of the window when tools return verbose output. - **Working state and output contract.** Required task state and response constraints. ## Agentic context engineering: why long runs are harder A single-turn application assembles context once. An agent assembles it at every step, which creates three additional constraints. 1. **The agent generates more context.** Tool calls, results and intermediate output accumulate during the run. 2. **Contradictions accumulate.** A value read at step 2 and updated at step 40 can remain in the same window without a marker for the current value. 3. **Cost grows with the square of the run.** Each step resends what came before, so a long run pays for its history repeatedly. Long-running agents therefore benefit from external storage that maintains current state and supplies selected information to each call. ## Four context operations Context management uses four operations. A system may combine them. ### Select Retrieve and rerank evidence relevant to the subject of the question. Apply user, tenant and project scopes before adding evidence to the prompt. ### Compress Summarize history, truncate tool output or extract the required passages from a document. Keep the source content available because compression removes detail. ### Isolate Run a subtask in a separate context window and return only the required result to the parent task. ### Persist Store data outside the context window and retrieve it when needed: files, a scratchpad, a database or memory infrastructure. This is the only option that keeps data available after the session ends. ## Where memory fits Selecting current information requires validity data in storage. Similarity ranking can give a superseded statement and its replacement similar scores. Without validity data, the model receives both statements without a reliable way to identify the current one. A memory system can track each fact's validity dates and return the current fact with dated evidence. In past.dev, `POST /api/v1/recall` returns ranked documents and source excerpts within a token budget. The application decides whether those documents establish an answer. [What is agent memory?](/what-is-agent-memory) explains the storage options. ## How to measure it Track four metrics for each version of the context assembler. - **Tokens per call, split by content type.** Measure instructions, tools, evidence, history and tool results separately. - **Cost per question.** The console dashboard records ingest and recall usage per call. - **Retrieval sufficiency**: how often the evidence needed to answer was actually in the window. Measured separately from whether the model then answered correctly, because they fail for different reasons. Our harness reports both, and the method is on [benchmarks](/benchmarks). - **Abstention rate.** Measure how often the system reports that evidence is insufficient. ## Frequently asked questions ### What is context engineering? The practice of deciding what goes into a model's context window on each call: instructions, tool definitions, retrieved evidence, conversation history, working state, and the output contract. Agent code assembles this context at runtime on every step. ### What is the difference between context engineering and prompt engineering? Prompt engineering optimises the wording of an instruction at design time. Context engineering decides the whole composition of the window at runtime: what is retrieved, what is compressed, what is isolated into a subtask and what is persisted outside the window. Prompt engineering is one component of it. ### Is context engineering just RAG? Retrieval is one context operation. Context engineering also covers compressing history, isolating subtasks in separate windows, persisting state outside the window, and allocating tokens among these operations. ### What is agentic context engineering? Agentic context engineering applies context selection to each step of a multi-step agent. Tool output and intermediate results increase context size. Contradictions can accumulate during the run. Cost increases when each step resends the previous history. ### How do I measure context engineering? Track tokens per call by content type, cost per question, retrieval sufficiency, answer accuracy, and insufficient-evidence rate. Evaluate cost and accuracy together because reducing context can remove required evidence. ## Related - [What is agent memory?](https://past.dev/what-is-agent-memory) - [Memory observability](https://past.dev/guides/memory-observability) - [Choosing a vector database for RAG](https://past.dev/vector-database-for-rag) - [Vector database vs graph database vs memory](https://past.dev/vector-database-vs-memory) - [Memory API overview](https://past.dev/docs/memory-api/overview) - [How we measure memory](https://past.dev/benchmarks)