--- title: "Why RAG gets less accurate as your data grows" description: "On BEAM at 10M tokens, the RAG in the BEAM paper scores 21.1% to 24.9% and past.dev 85.0%. Why RAG accuracy drops as your data grows." canonical: https://past.dev/blog/why-rag-gets-less-accurate date: 2026-10-09 category: Research authors: Louis Tissot --- # Why RAG gets less accurate as your data grows When I talk about past.dev, two questions come up every time: how it differs from RAG, and where RAG reaches its limits. I've answered them so often that I'm writing the answer down once, with numbers. By memory, I mean what an AI agent knows about your situation: what is true now, what changed and when, and where each fact came from. Without it, teams build agents for one narrow task each and write the rules by hand. They end up with deterministic workflows, and every new case needs a person to add it. The usual way to give an agent memory is RAG. But RAG gets less accurate as the history grows. BEAM is a public benchmark with conversations of up to 10M tokens. At 10M tokens, the RAG in the BEAM paper scores between 21.1% and 24.9%, depending on the model. past.dev scores 85.0%. ## How RAG searches a long history RAG searches your history for the text that looks most like the question. In the BEAM paper, each pair of user and assistant turns is one document, and the model gets the five documents closest to the question. This can work when the answer is in one turn. When a later turn changed the answer, or the answer is spread over several turns, the five documents can miss part of it. Say a customer tells your assistant in January that their budget is 10,000 euros, and in March that it went up to 15,000. When someone asks for the budget, both messages look like the question. RAG can return the old one, the new one or both, and the model has to guess. A longer history has more old messages that look like the question. The BEAM paper measures the effect: each of its four models scores lower at 10M tokens than at 100K. At 100K the paper's RAG scores between 26.9% and 32.3%, and at 10M between 21.1% and 24.9%. The authors write that models "struggle as dialogues lengthen", with and without retrieval. Some categories drop much more than others. Each cell gives the lowest and the highest score of the paper's four models. | What the question needs | BEAM category | At 100K tokens | At 10M tokens | |---|---|---|---| | One detail from one turn | Information extraction | 33.8% to 39.2% | 27.5% to 37.5% | | The newest value of a fact | Knowledge update | 27.5% to 37.5% | 30.0% to 37.5% | | Facts from several turns | Multi-hop reasoning | 14.8% to 26.3% | 5.0% to 12.5% | | A summary of the history | Summarization | 7.4% to 11.1% | 4.5% to 10.6% | | The time between two events | Temporal reasoning | 12.5% to 27.5% | 0.0% to 2.5% | | Two statements that conflict | Contradiction resolution | 1.8% to 5.0% | 0.0% to 2.5% | The best categories are a single detail and the newest value of a fact, and even there the paper's RAG scores at most 37.5% at 10M. Questions that need several turns, dates or two conflicting statements score 12.5% or less at 10M. You can't run customer support or an internal workflow on a system that scores 30%. ## What the usual fixes do The paper tested the three usual fixes: - Retrieve more documents: in a test of the paper's own method, LIGHT, the score rose as retrieval went from 5 to 15 documents, and fell at 20. The authors put this down to noisy context. - Change the search method: sparse, keyword-style retrieval instead of embedding search did not change the score much. - Send the full history: at 1M tokens, the two models with a 1M-token window score lower with the full history than with RAG. One scores 19.1% against 30.2%, and the other 19.9% against 27.1%. The full history is also expensive. Each question sends about 1M tokens to the model. At Claude Sonnet 5.5's list price of $2 per million input tokens, that is about $2 per question, or $2,000 per 1,000 questions. At 10M tokens it is not an option at all, because none of the four models in the paper can read 10M tokens. The paper's long-context baselines could only read the most recent part of each 10M conversation, and they score between 10.4% and 13.3%. ## How past.dev answers instead past.dev does most of its work when you send it data, before anyone asks a question. It reads each turn once and saves the facts in it as memories. Each memory records when the fact became true, when it was replaced, and which turn it came from. Recall then searches those memories rather than raw text. You can also ask a question as of a past date. I won't go into the details of how we build it. The difference shows in the categories of the first table: - When two statements conflict, past.dev records which one replaced the other. At 10M tokens it scores 92.5% on knowledge update and 73.1% on contradiction resolution. The paper's best RAG scores 37.5% and 2.5%. - Each memory links back to the turn it came from. Recall matches the question against the memory, so it can find a detail even when the original turn does not look like the question. At 10M tokens past.dev scores 68.8% on information extraction, against 37.5% for the paper's best RAG. With past.dev, the answering model gets at most 8,000 tokens of evidence per question, whether the history is 100K tokens or 10M. Cost matters to me as much as accuracy, and an answer that needs 1M tokens is too expensive to give every customer. The expensive step happens once, at ingestion, where a model reads each turn. ## Results The table compares past.dev with the best of the paper's four RAG runs, which is Llama 4 Maverick at every size. | History per conversation | past.dev | Best RAG in the paper | Difference | |---|---|---|---| | 100K tokens (20 conversations, 400 questions) | 92.1% | 32.3% | +59.8 points | | 500K tokens (35 conversations, 700 questions) | 89.6% | 33.0% | +56.6 points | | 1M tokens (35 conversations, 700 questions) | 90.7% | 30.7% | +60.0 points | | 10M tokens (10 conversations, 200 questions) | 85.0% | 24.9% | +60.1 points | Scores are mean BEAM scores in percent; the paper prints them from 0 to 1. The RAG numbers are from Table 1 of the paper. On four categories alone (contradiction resolution, information extraction, knowledge update and multi-hop reasoning, 80 questions at 10M), past.dev scores 65.6% and the paper's best RAG 20.6%, a gap of 45.0 points. The paper's four models score between 17.5% and 20.6% on them. Our harness calls multi-hop reasoning "multi-session reasoning". LIGHT's best results are 35.9% at 500K and 26.6% at 10M, which is 54 to 58 points below past.dev. At 10M tokens past.dev keeps 92% of its 100K score, and the paper's RAG keeps between 71% and 78%. I know a gap of almost 60 points looks too big to be true. Our [harness is public](https://github.com/pastdotdev/benchmarks), with every answer and every judge verdict, so you can check it yourself. ## How we tested past.dev's scores come from our public harness. The RAG scores are the ones the BEAM authors published; we did not run their RAG again. - Benchmark: [BEAM](https://arxiv.org/abs/2510.27246) (ICLR 2026), all four sizes, with 2,000 questions over 100 conversations in ten categories. - past.dev: the published run of 29 September 2026 on [the benchmarks page](/benchmarks), through the public API. The model openai/gpt-5.6-luna, with reasoning off, both answers the questions and judges the answers. It uses the answer and judge prompts listed in our harness. Each question gets up to 8,000 tokens of evidence. - The paper's RAG: each user and assistant turn pair is one document, embedded with BAAI/bge-small-en-v1.5 in a FAISS index, and the model gets the five closest documents. The answer models are Qwen 2.5 (32B), Llama 4 Maverick, Gemini 2.0 Flash and GPT-4.1-nano, with the paper's answer prompt and judge. - Scoring: the paper's method, in the version our harness uses. A judge scores each rubric item 0, 0.5 or 1, and a question scores the average. Event ordering uses Kendall tau-b. ## Try it on your own data You can [sign up](/) and test past.dev on your own data. If something here isn't clear, send me a message and I'll explain it again. ## Sources - [Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs](https://arxiv.org/html/2510.27246v1), Table 1 and section 4.1, for each RAG, LIGHT and long-context score, and section 4.2 for the retrieval tests - [BEAM code and data](https://github.com/mohammadtavakoli78/BEAM) - [past.dev's BEAM harness and results](https://github.com/pastdotdev/benchmarks/blob/main/beam/README.md), for each past.dev score by size and by category