--- title: "How we benchmark memory" description: "How past.dev benchmarks memory with corrected answer keys, strong baselines, multiple judges, and cost reported with accuracy." canonical: https://past.dev/blog/measuring-memory-honestly date: 2026-08-25 category: Research authors: The past.dev team --- # How we benchmark memory Memory benchmark results depend on the questions, answer keys, baselines and judges. Cost and latency also affect whether a system is suitable for production. This article describes the evaluation method used for past.dev. ## Required test cases A memory benchmark should include the following cases. - **Updates.** State a fact, revise it later and ask for the current value. - **Identity resolution.** Refer to the same person through an email signature, transcript name and chat handle. - **Historical questions.** Ask for the value that was valid at a specified date. - **Unsupported questions.** Ask questions that cannot be answered from the supplied history. These cases test functions beyond document retrieval. The [agentic memory article](/blog/agentic-memory) describes the underlying memory operations. ## Answer keys Published benchmarks can carry errors in their own answer keys. LoCoMo is the documented case: the community has corrected a portion of its keys, and a score differs materially depending on which set was used. Any published result should name the key version it ran against, so that a label error is not counted as a memory error. ## Baselines Every run includes two baselines. - **Full context** sends the complete history to the model for every question. - **Plain RAG** retrieves passages from the same corpus. These baselines show whether the memory system improves on simpler approaches. Run every baseline on the same questions and history size as the memory system. ## Retrieval and answer metrics Retrieval sufficiency checks whether the returned evidence contains all information required to answer the question. It also checks whether the current fact is ranked before superseded facts. Answer accuracy checks whether a model produced the expected answer from that evidence. This score depends on both retrieval and the selected answer model. We report the two measurements separately. ## LLM judges An LLM judge grades model output. Different judges can assign different scores to the same answer. Use multiple judges and report their results separately. This shows whether a result depends on one judge. Record each judge model, version, prompt and scoring rule with the run. ## Cost and latency Report accuracy, cost and latency from the same run. The production console records usage for each API call. [How to reproduce this](/benchmarks/methodology) ## Reporting requirements A published result should include enough information for another team to repeat the run. - Name the dataset, split and answer-key version. - Identify the tested system version and configuration. - State the history size and number of questions. - List the retrieval limits and token budgets. - Identify the answer model and every LLM judge. - Report failed, timed-out and degraded requests. - Publish accuracy, latency and total cost for the same run. Do not combine results from different configurations in one headline score. Keep retrieval and answer scores in separate columns. Include the baseline results in the same table. These requirements make configuration changes visible and allow readers to compare repeated runs. ## Evaluation on customer data The [benchmarks page](/benchmarks) reports current BEAM results and the broader evaluation scope. The [BEAM leaderboard](/benchmarks/beam) ranks them against every published result. Teams should also test their own data. Include changed sources, identity scoping, historical questions and unsupported questions. For unsupported questions, score whether the answer layer abstains instead of inventing support.