--- title: "What is agent memory?" description: "Agent memory stores information across sessions and retrieves relevant parts for later model calls. Learn the main memory types, architectures and limits." canonical: https://past.dev/what-is-agent-memory last-updated: 2026-10-10 --- # What is agent memory? Source: https://past.dev/what-is-agent-memory Agent memory stores information outside the model's context window and retrieves relevant information for later calls. Unlike a context window, it remains available after a session ends. It has four common functions: storing current task state, retaining facts across sessions, resolving entities across sources and tracking when facts change. ## What problems agent memory solves 1. **Continuity.** The user said something three sessions ago that changes the answer now. Without memory, every session starts from zero and the user repeats themselves. 2. **Working state.** Long tasks accumulate plans, intermediate results and decisions. External storage keeps this state available when it no longer fits in the context window. 3. **Identity.** The same person, account or project appears under different names in different sources. Something has to decide that they are the same subject, or the history stays split. 4. **Currency.** Facts change. The memory system must identify current and previous values. ## Agent memory vs a context window The context window contains the input for one model call. It has a fixed capacity, and its input tokens are charged on every call. Agent memory stores information outside that window and selects information for later calls. Larger windows can hold more history, but sending the full history on each call increases cost and adds irrelevant content. [Context engineering](/context-engineering) selects what to include. Memory supplies the stored information available for selection. ## The vocabulary: working, episodic, semantic, procedural The field borrows terms from cognitive psychology as labels for four storage functions. These labels do not describe model internals. | Term | What it holds | Typical implementation | | --- | --- | --- | | Working, or short-term | The current task: recent turns, intermediate results, the plan | The context window itself, plus a scratchpad the agent rewrites | | Episodic | What happened, and when: past conversations and events | A log of messages or events, retrieved by similarity or by time | | Semantic | Facts that hold independently of when they were said | Extracted statements in a store, a knowledge graph, or a profile document | | Procedural | How to do things: instructions, tools, learned preferences | The system prompt, tool definitions, and files the agent edits | Production systems often combine all four types. Compare products by the specific types and operations they implement. ## Agent memory vs RAG RAG retrieves from a corpus you curated to answer a question. Memory accumulates from the agent's own history and has to maintain it: merge duplicates, supersede outdated facts, and decide what is worth keeping. Both systems retrieve data and often use embeddings. The main difference is ingestion. RAG queries an existing corpus. Memory continuously adds and updates information from agent activity. ## Common agent memory architectures Five designs cover different storage and retrieval requirements. ### Buffers and summaries Keep the last N turns and summarise older turns. This is inexpensive, requires no additional infrastructure and is the default in most frameworks. Details omitted from a summary cannot be recovered from that summary. ### Vector recall over past messages Embed every message and retrieve similar messages at question time. This supports larger histories and requires no extraction step. It does not resolve contradictions: two conflicting messages can both match the question, and the index does not identify which one is current. ### Self-editing memory blocks The agent edits its own memory with tools. [Letta](https://docs.letta.com), which developed from MemGPT, keeps memory blocks in the context window and archival memory outside it. Correctness depends on when the agent chooses to write and what it records. ### Extracted facts in a store Extract statements from conversations, store them and retrieve them by user or session. [mem0](https://github.com/mem0ai/mem0) is an open-source implementation of this design. It uses a model during ingestion to select and write memories. ### Temporal knowledge graphs Extract entities and relationships and give each fact a validity window. Superseded facts remain available for historical queries. [Graphiti](https://github.com/getzep/graphiti), the open-source engine behind Zep, is one documented example. It requires extraction, entity resolution and update handling during ingestion. [Knowledge graphs for LLM applications](/knowledge-graph-for-llm) covers the tradeoffs. ## Agent memory requirements - **Identity resolution.** References to the same person or organization must resolve to one entity, including when new identifiers arrive later. - **Changed facts.** The system must store when a fact was valid and retain the previous value after an update. - **Retention.** Define which records are retained, compressed, or deleted. Test retrieval as the stored history grows. - **Insufficient evidence.** The response must report when stored evidence does not support an answer. - **Cost.** Extraction and retrieval both add cost. Track usage per ingest and recall operation. ## How to tell if your agent needs memory - Users repeat context they already gave you in a previous session. - The agent answers with something that was true last quarter. - The same subject appears under different names and the agent treats them as strangers. - Your prompt has grown a section that pastes history in, and it keeps growing. - You cannot identify the source of an answer. If none of these conditions applies, a retrieval pipeline may be sufficient. See [choosing a vector database](/vector-database-for-rag). ## Frequently asked questions ### What is agent memory? Agent memory is external storage and retrieval logic for information from previous sessions. It stores facts, events, preferences, and task state, then selects relevant information for a later model call. ### What is the difference between agent memory and a context window? A context window contains the input for one model call and has a fixed token capacity. Agent memory persists outside the model and supplies selected information to later calls. ### Is RAG the same as agent memory? No. RAG retrieves from an existing corpus. Memory also ingests and maintains the agent's history by merging duplicates, superseding outdated facts, and resolving identities. Retrieval may use similar methods in both systems. ### Do I still need memory with a one million token context window? Yes for long-running agents. Sending the full history on every call increases input cost and can reduce answer accuracy when irrelevant or conflicting material competes for attention. A larger window adds capacity but does not select the required information. ### What is the difference between short-term and long-term memory in an agent? Short-term, or working, memory is the current task's state, usually the context window plus a scratchpad. Long-term memory remains available after the session: episodic records of what happened, semantic facts about the world, and procedural instructions about how to act. ### How do I store agent memory? Four common options: keep raw messages and summarise, embed messages and retrieve by similarity, extract facts into a store, or build a temporal graph where each fact carries a validity window. They differ mainly in how much work happens at write time, and in whether contradiction can be represented at all. ## Related - [How to choose a memory system](https://past.dev/guides/choose-memory-system) - [Facts that change over time](https://past.dev/guides/facts-that-change-over-time) - [Context engineering for AI agents](https://past.dev/context-engineering) - [Vector database vs graph database vs memory](https://past.dev/vector-database-vs-memory) - [Knowledge graphs for LLM applications](https://past.dev/knowledge-graph-for-llm) - [Memory API overview](https://past.dev/docs/memory-api/overview) - [How we measure memory](https://past.dev/benchmarks)