--- title: "You'll never reach 85% on BEAM10M using .md files" description: "Why agent memory kept in .md files falls short on BEAM at 10M tokens, and how the past.dev Memory API handles changing facts, late records and permissions." canonical: https://past.dev/blog/md-files-vs-memory-api-beam-10m date: 2026-09-22 category: Engineering authors: The past.dev team --- # You'll never reach 85% on BEAM10M using .md files The most common way to give an agent memory is a Markdown file. [Claude Code](/blog/claude-code-memory), [OpenClaw](/blog/openclaw-memory) and [LLM wikis](/blog/llm-wiki) all work this way. The agent reads the files into its context and edits them as it learns. BEAM, the long-term memory benchmark, asks questions over conversation histories of up to 10M tokens. At 10M, the past.dev Memory API reaches 84.95%, the highest published result. The next is 68.0% ([leaderboard](/benchmarks/beam)). A `.md` file will not get there. We learned why from agents we ran for teams: they followed a team's meetings, email, chat and knowledge base, and answered questions about them. The past.dev Memory API grew out of those agents. Take a product launch. Michael owns it from January, and in March a private email hands it to Andy. Then in April someone imports an old mailbox, which shows that Dwight owned the launch for a few weeks in February. Keep this in a notes file, and an agent asked "who owns the launch?" runs into four problems: 1. **The facts changed.** If the agent rewrites the line to say Andy, the file loses that Michael owned the launch in January. If it appends, the file names two owners and the agent has to guess which one is current. 2. **A record arrived late.** Dwight's February assignment is written in April, after Andy's. The file puts Dwight last, and the agent reads Dwight as the current owner. 3. **A source is private.** Every agent that loads the file reads all of it. Once the move to Andy is written down, anyone whose agent loads the file learns about the private email. 4. **The file grows.** Reading the whole history costs more with every call, and in the BEAM paper, models given the full history score at most 25.9% at 1M tokens. So the agent summarizes the file, which drops details, or searches it, which returns passages without dates or relations. The Memory API works out what is current, the order in which things happened, and who may see what once, when data comes in. **TL;DR** - `.md` file memory loses history when it is rewritten, misorders late records, shows every reader everything, and has to be summarized or searched as it grows. - Memory stores facts and how they relate. Nothing is overwritten, so any date can be queried. - Entity recognition, inference models and LLM judges decide what facts are about and how they relate. If the evidence is unclear, the facts stay unlinked. - Every fact keeps its sources' permissions, and search applies them before ranking. - Recall ranks by what the question is asking, such as the current state or the full history. - On BEAM at 10M tokens, past.dev scores 84.95%, the highest published result. ![The Memory API at a glance: four sources about the launch, one of them private, are linked once at ingestion into one Launch entity and three dated owners, Michael, Dwight and Andy, each replacing the one before. A question and the person asking enter at Andy, and recall returns a page with Andy now and the earlier owners dated.](/blog/engineering-memory/overview.svg) ## Temporal memory: what memory records The first two problems are about time. A `.md` file keeps text in the order it was written. To handle time, memory has to store facts, their dates and how they relate. Each source is broken into **memory entities**: facts, the people and things they are about, and the links between them. Every memory entity points back to the passage it came from. When a new fact arrives, memory links it to earlier ones: a follow-up confirming Michael **corroborates** his assignment, and the move to Andy **supersedes** it. ![Chunks versus memory entities: three dated passages about the launch owner, two in conflict, leave the agent to infer which is current; the same sources as memory entities give two dated facts linked to Michael, Andy and the launch, with a supersedes link and the March fact marked current.](/blog/engineering-memory/entities.svg) A replaced fact stays in memory with the date it was replaced, so "Who owned the launch in January?" still returns Michael while "Who owns it?" returns Andy. Facts are ordered by when they happened, whatever order they arrive in. So Dwight's assignment, imported in April, lands in February, between Michael and Andy. Andy stays the current owner. ![Late evidence fills a gap in the ownership history. A February assignment to Dwight arrives in April and is placed between Michael in January and Andy in March.](/blog/engineering-memory/history.svg) These links decide what recall returns, so a wrong link hides a true fact from every later search. ## Continuous consolidation: getting the links right Memory has to decide that the move to Andy replaces Michael's assignment. Matching on wording is not enough. "Let's give the launch to Andy" looks like "Andy now owns the launch", but only one of them is a decision. A handoff on a different launch looks the same too. Memory first works out who and what each fact is about. A **named entity recognition (NER)** model finds the people, projects and other things a source mentions. **Entity resolution (ER)** then decides which mentions are the same thing. "Michael Scott" in a meeting, "michael" in chat and michael@office.com resolve to one person, and "the Q3 launch" and ticket LAUNCH-42 to one project. ![Entity resolution across sources: a full name in a meeting, a first name in chat and an email address resolve to one person, and a thread's "the Q3 launch" and a ticket number resolve to one project. The bare first name stays provisional until other sources agree.](/blog/engineering-memory/resolution.svg) Then it decides how facts about the same entities relate: - **Natural language inference (NLI)** checks whether one fact supports or contradicts the other. - **LLM judges** take the cases the model cannot settle. No LLM answer is trusted alone. Judges see the source passages behind each fact, so every verdict, for or against a link, is made against the original passages. A suggested identity or link stays provisional until other sources agree. Contested cases go through several adversarial passes. If doubt remains, the facts stay unlinked until later evidence settles the case. This process runs continuously as data arrives. Probes also revisit older memory, so stored facts gain support or get replaced as new evidence comes in. ## Permission-aware retrieval: who can see what The email about Andy is private. Anything written in a `.md` file is visible to every agent that loads it. A team agent answers many people, so memory checks who is asking before it returns the email or anything built from it. The developer sends permissions along with the data. Each source comes with an **audience**: the people allowed to read it. Each recall names the person the agent acts for. Facts inherit their sources' audiences. A fact built from a shared doc and a private email is visible only to people who can see both. Links follow the same rule. Showing that Michael's assignment was replaced would reveal the private email. So people who cannot see the email simply see Michael as owner. Search applies audiences before ranking. Filtering afterwards fails people with narrow access: if nineteen of the top twenty results are private, only one is left. ![Illustrative top-k comparison: global selection contains one readable result and nineteen inaccessible results. Filtering by audience before selection reserves all candidate slots for readable evidence, when enough matches exist.](/blog/engineering-memory/audiences.svg) Permissions work because every fact remembers its sources. The same record lets new evidence strengthen a fact instead of duplicating it. It also makes deletion complete, including for right-to-be-forgotten requests: deleting a source also deletes every fact that depended on it alone. ![Lineage, accumulation and deletion: three sources feed one fact and one of them alone feeds a second derived item; after that source is deleted, the item derived only from it is gone and the fact stands on the two remaining sources.](/blog/engineering-memory/lineage.svg) ## Answering the question With facts linked and permissions attached, the last step is answering. "Who owns the launch now?" and "How did ownership change?" need the same facts, in a different order. Recall starts with standard hybrid search: vector and keyword matching, query expansion, rank fusion and a reranker. Then it uses the links built earlier. A search hit on Andy's handoff pulls in Michael's earlier assignment, the notes confirming it, and the people involved, even if they share no words with the question. ![Search hits as anchors into the graph: the hit that Andy owns the launch pulls in the fact it superseded, the notes that corroborate it, and the project and person it concerns, none of which matched the question directly.](/blog/engineering-memory/anchors.svg) Recall also reads what the question is after: a fact, the current state, a sequence, a list, an explanation or a preference. Asked "Who owns the launch now?", recall puts Andy first and lists Michael and Dwight after him with the dates their ownership ended; asked "How did ownership change?", it returns all three in the order they held it. Cognitive neuroscience describes the same behavior in people as [retrieval orientation](https://pubmed.ncbi.nlm.nih.gov/10689345/): we search memory differently depending on what we are trying to recall (Rugg and Wilding, 2000). ![What recall stacks on top of retrieval: the standard toolkit at the base, the knowledge graph consolidation built above it, intent and time above that, and the page with lineage on every result at the top.](/blog/engineering-memory/recall.svg) The result is a short, deduplicated page within a token budget. Every item shows its dates and the exact source passages. ## Cost compared with a .md file An agent with a `.md` file sends the file, or the history behind it, with every call. Each call costs more as the history grows. Fortunately, it can only grow as far as the context window, but you miss more and more of the facts through compaction. Our post on [cutting LLM token costs](/blog/llm-token-costs) works through the numbers. With memory, history is written once. The agent sends its question and the person it is asking for, and memory returns a page of evidence: 8,000 tokens by default, up to 32,000, however large the archive. On our benchmark workload, that costs 20x less per question than sending 1M tokens of history. ![Tokens per question: re-sending history grows with every call until it crosses the context window, while recall from memory stays at the evidence budget as the archive grows from 100K to 10M tokens.](/blog/engineering-memory/context.svg) The archive can keep growing without the developer pruning a `MEMORY.md`, writing summaries, filtering per user or defining a schema. ## Results and limitations On BEAM, the long-term memory benchmark, past.dev ranks first at every history size among published results, from 100K to 10M tokens. At 10M, it scores 84.95%, while the next published result is 68.0%. The [BEAM leaderboard](/benchmarks/beam) has every figure with its source, and the [methodology](/benchmarks#method) shows how to reproduce ours. Known limits: - Questions that combine evidence from unrelated conversations lose the most accuracy at scale. - No published BEAM result measures a `.md` file memory directly. The comparison in this post rests on the BEAM paper's full-history runs. - BEAM measures answers only. Privacy, deletion and recovery need their own tests. - Entity types are general-purpose. Very specific domains may need a custom ontology and a custom extraction model, e.g. [GLiNER2](https://github.com/fastino-ai/GLiNER2). We build those with Enterprise customers and design partners. ## Getting started The [Memory API quickstart](/docs/memory-api/quickstart) covers ingestion and recall. The [benchmarks page](/benchmarks) has the current figures and configurations.