--- title: "Knowledge graph for LLM applications" description: "How knowledge graphs support multi-hop questions, provenance and entity resolution in LLM applications, including build costs and temporal data." canonical: https://past.dev/knowledge-graph-for-llm last-updated: 2026-10-10 --- # Knowledge graph for LLM applications Source: https://past.dev/knowledge-graph-for-llm A knowledge graph stores typed entities and explicit relationships. An LLM application can traverse those relationships to answer multi-hop questions and return the path as evidence. Temporal queries require additional fields for when each fact became and stopped being valid. Those validity periods let the graph distinguish current and previous facts. ## What is a knowledge graph in an LLM application? A knowledge graph represents entities as nodes and typed relationships as edges. An LLM application converts a question into a graph query and adds the returned subgraph to the model context. A vector index stores passages and returns similar passages. A knowledge graph stores entities and relationships extracted from source material. A graph traversal can combine facts from several sources. ## What a knowledge graph does well - **Multi-hop questions.** For example, identify who reports to the person who owns the account that raised a ticket. The required relationships may be distributed across several sources. - **Provenance.** Each edge can reference the source from which it was extracted. - **Deduplication through identity.** Once two names resolve to one node, facts associated with either name refer to the same subject. - **Constraints.** A typed schema can reject relationships that violate the schema before they reach a reader. - **Explainability.** The application can show the path used to produce an answer. ## What it costs to build Structured records already contain entities and keys. Building a graph from prose also requires extraction, entity resolution, schema management and update handling. 1. **Extraction.** A model reads each document and proposes entities and relationships. This puts a language model in the write path, so ingestion cost scales with corpus size and the output is not deterministic. 2. **Entity resolution.** Each proposed entity must be matched against existing entities to prevent near-duplicate nodes. New evidence can require previous identity decisions to be updated. 3. **Schema.** A fixed ontology limits the available types. An inferred ontology can produce inconsistent types that require normalization. 4. **Maintenance.** New documents can contradict old edges. The system needs a defined update policy. A small, static corpus may need only [vector or keyword retrieval](/vector-database-for-rag). ## Adding time to a knowledge graph A basic graph records whether an edge exists. Overwriting an edge removes the previous value. Adding a second edge preserves both values but does not identify which one is current. A temporal graph records when each fact became valid and when it stopped being valid. It retains superseded facts for historical queries and selects open validity windows for current-state queries. A bitemporal model records when the fact was true and when the system learned it. The dates differ for backfilled or corrected data. Keeping both prevents information learned later from appearing in earlier knowledge-time queries. ## Graphiti and Zep [Graphiti](https://github.com/getzep/graphiti) is an Apache 2.0 framework for temporal graphs maintained by Zep. It models ingested data as episodes and derives entity nodes and relationship edges. Each fact has a validity window and a link to its source episode. New information can invalidate an earlier fact while retaining it for history. Graphiti runs on Neo4j, FalkorDB, and Amazon Neptune. Zep is the managed product built on Graphiti. The team also published the architecture as a paper on temporal knowledge graphs for agent memory. Graphiti is available to teams that want to operate and customize the graph engine themselves. past.dev ingestion accepts raw text with its original timestamp without a required schema. `/recall` returns ranked documents with occurrence dates and source excerpts, and the application decides whether they establish an answer. See [past.dev vs Zep](/vs/zep) for the detailed public-contract comparison. ## Do you need to design an ontology? A domain graph over structured records often benefits from a defined ontology. A graph extracted from conversations and documents can infer types during ingestion, then normalize them as more data arrives. The past.dev ingest contract does not require an ontology. Applications send text and an optional timestamp, label, source id, metadata, audience slug, and author identity. The API does not expose graph or ontology controls. ## Frequently asked questions ### What is a knowledge graph for an LLM? A store of typed entities and explicit relationships that an LLM application queries by traversal instead of by similarity. It answers multi-hop questions, carries provenance on every edge, and makes an answer explainable as a path. ### Knowledge graph vs vector database: which is better for RAG? A graph supports precise multi-hop queries over modeled entities and can combine facts from several documents. A vector index covers all embedded passages and supports open-ended similarity queries. Many systems use both. ### Can an LLM build a knowledge graph automatically? It can propose entities and relationships from text. The system must then match each proposal against existing entities and define how new information updates or conflicts with stored edges. ### What is a temporal knowledge graph? A knowledge graph where each fact carries a validity window, so a superseded fact is marked invalid from a date rather than deleted. It answers what is true now and what was true then from the same store. Bi-temporal graphs keep two clocks: when a fact held in the world, and when the system learned it. ### Do I need Neo4j to use a knowledge graph with an LLM? No. A dedicated graph database is one option, and Graphiti supports Neo4j, FalkorDB and Amazon Neptune. Other products expose the behavior through a managed API, where the underlying storage is not part of the integration contract. ## Related - [past.dev vs Zep](https://past.dev/vs/zep) - [Vector database vs graph database vs memory](https://past.dev/vector-database-vs-memory) - [What is agent memory?](https://past.dev/what-is-agent-memory) - [How the pipeline works](https://past.dev/docs/memory-api/how-it-works) - [How we measure memory](https://past.dev/benchmarks)