--- title: "How AI agents read a website" description: "How AI agents read a website through HTML, Markdown and llms.txt interfaces, with tests for every published format." canonical: https://past.dev/blog/website-agents-can-read date: 2026-08-26 category: Engineering authors: The past.dev team --- # How AI agents read a website We maintain past.dev as a website AI agents read through HTML and Markdown interfaces. AI agents account for more visits to past.dev than human users. Each page is published as HTML for browsers and Markdown for agents. Both formats use the same content source. This article describes the published interfaces and the checks used to keep them consistent. ## Audit results Is Agentic gives past.dev a score of 100/100. Ora gives the site 80/100, grade B. Its remaining deductions concern external package registries and search presence. Lighthouse gives every page 100 in its agentic-browsing category. The audits use different requirements. We run all three and review each reported failure. ## How AI agents read the website Documentation pages use a typed content module. The HTML page and Markdown version render from that module. Blog Markdown files are served directly as their Markdown versions. We added an automated consistency test after a documentation page and its Markdown version contained different copy. The build now fails if either format uses separate prose. ## llms.txt and other machine-readable interfaces - [/llms.txt](/llms.txt) lists the available pages and links to their Markdown versions. - `/docs/llms-full.txt` contains the full documentation in one file of about 15,000 tokens. - `Accept: text/markdown` returns the Markdown version of a requested page. - Known AI crawlers receive Markdown automatically. - Machine-readable URLs keep a file extension and are exempt from rate limiting. - The `.well-known` directory contains the ARD catalog, RFC 9727 API catalog and agent-skills index. - `robots.txt` permits named AI crawlers. Its Content Signals policy allows search and AI input and disallows AI training. ## CSS delivery CSS is served as a linked file. Inlining it would increase the homepage response from 80 KB to 165 KB. It would also move the first documentation heading to byte 77,000. With linked CSS, the first heading appears within the first 19 KB. This arrangement allows HTML parsers to reach the page content earlier. The mobile Lighthouse performance score is about one point lower. ## Text requirements Important information must appear in text. Do not communicate status through color or position alone. Show API identifiers exactly as defined. Place evidence next to claims that an agent may quote. The same requirements apply to our [benchmark reports](/blog/measuring-memory-honestly). ## Build validation The build checks every registered page. Each page must have one heading, a canonical URL and a non-empty Markdown version. Documentation links must resolve to registered routes. Markdown code fences and tables must be valid. The generated `llms.txt`, full documentation file, sitemap and RSS feed must include the expected entries. The tests also scan published copy. They reject old brand names, unsupported API fields, em and en dashes, and common marketing phrases. A failed check stops the production build. ## Implementation checklist 1. Publish an `llms.txt` index. 2. Generate HTML and Markdown from the same source. 3. Provide the complete documentation as one file. 4. Keep machine-readable URLs available to automated clients. 5. Test every output format during the build. 6. Run more than one agent-readiness audit. For product documentation, start with the [Memory API quickstart](/docs/memory-api/quickstart). [What is agent memory?](/what-is-agent-memory) explains the product category.