--- title: "Introducing past.dev: memory for AI agents" description: "past.dev, a memory layer for AI agents, is available today. An agent asks it for context and gets current facts with dated sources, limited to what it may see." canonical: https://past.dev/blog/introducing-past-dev last-updated: 2026-09-29 --- # Introducing past.dev Announcements, September 30, 2026. Source: https://past.dev/blog/introducing-past-dev past.dev, a memory layer for AI agents, is available today, and sign-up is self-serve. It gives an agent what a longer context window does not: a record of what is true now, what it replaced, and when each fact changed. Your application sends past.dev the text it already produces, emails, call notes, tickets and chats, each stamped with the time it happened. When an agent needs context, it asks past.dev and gets back the current facts with the dated sources behind them, limited to what that reader is allowed to see. past.dev is state of the art on BEAM, the benchmark for long-term memory in agents, at every history size where memory systems have published a result: **BEAM accuracy by history size** | System | 100K | 500K | 1M | 10M | | --- | ---: | ---: | ---: | ---: | | past.dev | 92.08% | 89.63% | 90.65% | 85.03% | | Exabase M-1 | 76.9% | not published | 75.0% | 68.0% | | Hindsight | 75.0% | 71.1% | 73.9% | 64.1% | | Honcho | 63.0% | 64.9% | 63.1% | 40.6% | | mem0 | not published | not published | 64.1% | 48.6% | Answer accuracy on BEAM, complete splits, higher is better. Each competitor figure is that vendor's own published result; a missing bar means no published figure at that size.[^1] The Memory API is open to teams building agents today: [get an API key](https://sso.past.dev/sign-up), self-serve, with no card required. The [Pay as you go plan](https://past.dev/pricing) includes 150,000 credits ($45) to start, or 400,000 credits ($120) if you sign up with a work email, and Flex starts at $99 a month with unlimited recall calls. Storage and seats are never metered. Early-stage startups can apply for [six months of Flex at no charge](https://past.dev/startup-program). ## Why agents forget Most agents remember in one of two ways. They resend the conversation on every call, which gets slower and more expensive as the history grows. Or they search a vector index, which returns the text that looks most like the question. Neither knows which statement is current. Take an account team running a pilot with Acme. On 3 February, a call note records a budget of $32k. On 28 July, an email from the customer moves it to $40k. Both records are true, and both mention the budget. A similarity search ranks them by wording rather than by time, and the agent answers with whichever reads better. Asked the same question, past.dev returns $40k, cites the email of 28 July, and keeps the call note of 3 February as the value it replaced. The problem grows with the history. In the [BEAM paper](https://arxiv.org/abs/2510.27246), a model given the whole conversation and no memory system scores under 30% at 128K tokens and under 14% at 10M.[^2] ## How it works past.dev is three calls to the Memory API. 1. **Send what happened.** `POST /api/v1/ingest` with the text, the time it occurred and its audience. Emails, transcripts, tickets and documents go in as they are, and past.dev derives the facts, rules, names and dated events in them. 2. **Wait until it is readable.** Ingestion runs in the background, and `GET /api/v1/ingest/{ingestionId}` reports when it has completed. 3. **Ask as a reader.** `POST /api/v1/recall` with a question and an identity. The response lists ranked documents, each with its date, the verbatim excerpts behind it and the ids of your source records. Add `queryTimestamp` to read memory as it stood on a past date. ```http POST /api/v1/recall { "query": "What is the Acme pilot budget?", "identity": "demo-user" } ``` First result of the response, abridged: ```json { "asOf": "2026-09-22T10:00:00Z", "results": [{ "rank": 1, "occurredAt": "2026-07-28T16:00:00Z", "content": "Acme pilot budget: currently 40k.", "artifact": { "kind": "claim" }, "sources": [{ "sourceId": "call-8821", "occurredAt": "2026-07-28T16:00:00Z", "excerpts": ["Budget for the Acme pilot moved to 40k."] }] }] } ``` ## Evaluating past.dev ### BEAM BEAM, [published at ICLR 2026](https://arxiv.org/abs/2510.27246), asks 2,000 questions about 100 long conversations at four lengths, from 100K to 10M tokens, and grades ten abilities, from event ordering to contradiction resolution. We ran every split in full. The figures are from the run of 29 September 2026.[^1] past.dev answers 92.08% of questions at 100K tokens and 85.03% at 10M. Across a hundredfold increase in history it loses 7.05 points, and the next published system at 10M scores 68.0%. The breakdown by ability shows where memory still fails, and we publish it in full: **past.dev by BEAM ability** | Ability | 100K | 500K | 1M | 10M | | --- | ---: | ---: | ---: | ---: | | Event ordering | 100.00% | 100.00% | 99.81% | 100.00% | | Preference following | 100.00% | 99.05% | 98.87% | 100.00% | | Summarization | 99.38% | 97.25% | 99.60% | 99.00% | | Abstention | 98.75% | 92.14% | 97.14% | 100.00% | | Instruction following | 96.25% | 92.02% | 96.43% | 95.00% | | Temporal reasoning | 95.00% | 93.45% | 96.79% | 93.75% | | Knowledge update | 89.38% | 86.07% | 85.71% | 92.50% | | Information extraction | 95.00% | 86.32% | 83.98% | 68.75% | | Contradiction resolution | 78.13% | 78.57% | 75.36% | 73.13% | | Multi-session reasoning | 68.95% | 71.39% | 72.82% | 28.17% | past.dev accuracy by ability, run of 29 September 2026. 40 questions per ability at 100K, 70 at 500K and 1M, 20 at 10M. Scores under 80% are shaded. Six abilities stay above 90% at every history size. Four do not. Contradiction resolution sits between 73% and 79% at every size, and knowledge update between 86% and 93%. Information extraction falls to 68.75% at 10M tokens. Multi-session reasoning, which asks for facts stated in separate conversations to be combined, falls to 28.17% at 10M. These four are the work in front of us. ### LoCoMo LoCoMo is the long-conversation benchmark most memory systems report on. past.dev answers 93.12% of its 1,540 questions using the community-corrected answer key, with a macro average of 89.18% across its four categories.[^3] **LoCoMo accuracy** | System | Accuracy | Source | | --- | ---: | --- | | past.dev (corrected answer key) | 93.12% | [past.dev](https://past.dev/benchmarks) | | mem0 | 92.5% | [mem0.ai/research](https://mem0.ai/research) | | Hindsight | 92.0% | [benchmarks.hindsight.vectorize.io](https://benchmarks.hindsight.vectorize.io/) | | Honcho | 89.9% | [plasticlabs.ai](https://plasticlabs.ai/blog/research/Benchmarking-Honcho) | Accuracy on LoCoMo. past.dev uses the community-corrected answer key; the other three pages do not name the key they used, so read this comparison as indicative.[^3] ### Cost Answering from memory is also cheaper than rereading. past.dev is 20x cheaper per question than sending 1M tokens of history[^4], because the agent receives the facts it needs rather than the whole record. ## Memory you can check A memory layer is only useful if you can verify what it says. Every recall result carries the verbatim excerpt, its date and the id of the source record it came from, so an answer can be traced to the email or ticket behind it. When a new record changes a fact, the earlier value stays readable as history, and a recall with a past date reads memory as it stood then. Every record belongs to an audience, and every recall is made as one reader, so a support agent and a sales agent asking the same question each get what they are allowed to see. Deleting a record removes the memory derived from it. Customer data is never used to train models, and past.dev is [SOC 2 Type II audited and ISO 27001 and ISO 27701 certified](https://past.dev/security). ## What is in the API today - ingestion of one record or a batch, with its status; - recall as a reader, with `queryTimestamp` to read the past and `occurredFrom` and `occurredTo` for a period; - audiences as fixed lists of identities or rules over identity traits; - deletion by source record or by ingestion, which removes the derived memory; - the managed service on every paid plan, and a dedicated region or self-hosting on [Enterprise](https://past.dev/enterprise). ## Get started [Get API key](https://sso.past.dev/sign-up), and the [quickstart](https://past.dev/docs/memory-api/quickstart) takes you from that key to a first recall. The full comparison and the method are on the [benchmarks page](https://past.dev/benchmarks), and every published BEAM result is ranked on the [BEAM leaderboard](https://past.dev/benchmarks/beam). ## Footnotes [^1]: BEAM: complete splits at every size, 400 questions at 100K, 700 at 500K, 700 at 1M and 200 at 10M, in the run of 29 September 2026. Competitor figures are each vendor's own published result, linked from the [benchmarks page](https://past.dev/benchmarks). They are separately published runs, so models, judges and configurations differ between systems. Method: [How to benchmark agent memory](https://past.dev/benchmarks/methodology). [^2]: The [BEAM paper](https://arxiv.org/abs/2510.27246)'s no-memory baseline, measured on Qwen 2.5, Llama 4 Maverick, Gemini 2 Flash and GPT-4.1 nano. [^3]: LoCoMo: the locomo10-corrected split, with the answer key corrected by a public community audit and the adversarial category excluded, 1,540 questions across 10 conversations, run of 16 September 2026. The figure is the overall accuracy; the macro average is 89.18%. The other three figures are each vendor's own published result, and none of the three pages states which answer key it used: mem0, [mem0.ai/research](https://mem0.ai/research); Hindsight, [benchmarks.hindsight.vectorize.io](https://benchmarks.hindsight.vectorize.io/); Honcho, [plasticlabs.ai](https://plasticlabs.ai/blog/research/Benchmarking-Honcho). [^4]: Measured by the past.dev team at 1M tokens of history. ## Related - [Benchmarks](https://past.dev/benchmarks): The full comparison and the method. - [The BEAM leaderboard](https://past.dev/benchmarks/beam): Every published BEAM result, ranked at each size. - [How we benchmark memory](https://past.dev/blog/measuring-memory-honestly): Corrected answer keys, baselines, judges, and cost reported with accuracy.