--- title: "Vector database for RAG: how to choose one" description: "Compare pgvector, Pinecone, Qdrant and Weaviate for RAG by deployment, filtering, scale and hybrid search, with common retrieval failure modes." canonical: https://past.dev/vector-database-for-rag last-updated: 2026-10-10 --- # Vector database for RAG: how to choose one Source: https://past.dev/vector-database-for-rag Choose a vector database by operational requirements. pgvector keeps vectors in an existing Postgres deployment. Pinecone provides a managed index. Qdrant provides open-source self-hosting and payload filtering. Weaviate combines keyword and vector search in one query. These systems retrieve similar passages. Applications need separate logic to determine whether a retrieved statement is current. ## What a vector database does in a RAG pipeline A vector database returns stored embeddings nearest to a query embedding, optionally under metadata filters. The RAG pipeline separately handles chunking, embedding generation, reranking, prompt assembly and answer generation. Chunking, embedding and reranking choices can change answer quality even when the underlying vector index is the same. ## Vector database options for RAG The following six products cover common managed, self-hosted, embedded and distributed deployments. | Tool | Licence | How you run it | Use when | | --- | --- | --- | --- | | [pgvector](https://github.com/pgvector/pgvector) | PostgreSQL licence | An extension in your own Postgres | You already run Postgres and want vectors in the same transaction and backup | | [Pinecone](https://www.pinecone.io/) | Commercial, closed source | Fully managed only, no self-host | You want an index with no servers to operate and accept that data leaves your network | | [Qdrant](https://qdrant.tech/) | Apache 2.0 | Self-host, or their managed cloud | You want open source, heavy metadata filtering and control over quantization | | [Weaviate](https://weaviate.io/) | BSD 3-Clause | Self-host, or their managed cloud | You want BM25 and vector search fused in a single query, with vectorizer modules built in | | [Milvus](https://milvus.io/) | Apache 2.0 | Self-host, distributed, or managed | The index requires distributed storage and search | | [Chroma](https://www.trychroma.com/) | Apache 2.0 | Embedded, or a server | You need an embedded index for a prototype | ### pgvector, the default when you already run Postgres pgvector adds a `vector` type to Postgres with exact and approximate nearest-neighbour search. It supports HNSW and IVFFlat indexes. Its documentation states that vectors can contain up to 16,000 dimensions. Indexes support up to 2,000 dimensions for `vector` and 4,000 for `halfvec`. Rows and vectors can share a transaction, backup and access-control model. Applications can join embeddings to their source records in SQL. A dedicated vector engine provides more search-specific tuning and can scale independently from the relational workload. ### Pinecone, the managed default Pinecone is a hosted vector database with serverless indexes. Customers do not manage nodes, index rebuilds or capacity. It is closed source and has no self-hosted deployment, so evaluate its data location and procurement terms. ### Qdrant, open source with strong filtering Qdrant is written in Rust and released under Apache 2.0. It can run in a container or as a managed cloud service. Payload filters run during search. Quantization options reduce memory use with a corresponding change in recall. ### Weaviate, hybrid search in one server Weaviate is written in Go and released under BSD 3-Clause. A query can combine BM25 keyword scoring with vector similarity. Keyword scoring helps with exact identifiers, error codes and product names. Vectorizer modules can generate embeddings during insertion. ## Five selection questions 1. **Do you already run Postgres?** Start with pgvector if it meets the index-size, filtering, and latency requirements. A separate datastore adds backup, security, monitoring, and on-call work. 2. **Does every query filter by tenant, user or project?** Check whether filters run during nearest-neighbour search or after it. Post-filtering can return fewer results than requested because matching records may be excluded from the initial approximate result set. 3. **How large will the index get?** Millions of vectors fit on one node. Hundreds of millions usually require a distributed system. 4. **Do your queries contain exact tokens?** Part numbers, error codes, and names often require keyword scoring alongside vectors, either in the database or in a reranker. 5. **Who operates it?** Small teams may prefer a managed service to reduce on-call and maintenance work. ## Common RAG failure modes Many retrieval failures originate in ingestion, identity handling, ranking or response logic. - **Chunking splits a fact from its subject.** A chunk may contain 40k while the pilot name appears only in the preceding paragraph. The retrieved passage then lacks the subject required to use the value. - **Embeddings match related text.** A question about one budget can retrieve several passages about budgets without identifying the required value. - **Identity is not resolved.** The same person appears as a display name in a transcript, an email address in a thread and a handle in chat, so their history is three disconnected sets of chunks. - **Similarity does not determine validity.** February's number and July's number can be equally similar to the question. - **Top-k retrieval always returns results.** The application must determine whether those results contain enough evidence to answer. Reranking can improve chunk and similarity results. Identity resolution, temporal validity, and evidence sufficiency require additional data and logic. [Vector database vs graph database vs memory](/vector-database-vs-memory) compares these functions. ## Static documents and changing history Documentation, product pages and manuals are often updated in place. A vector database is appropriate when the application only needs the latest document version. Email, meeting transcripts, tickets and CRM notes are append-only histories. Current-state questions over these records require dates and update handling. past.dev accepts the text with its original timestamp and returns current information with dated evidence. ```bash curl -X POST https://api.past.dev/api/v1/recall \ -H "Authorization: Bearer $PAST_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "query": "what is the Acme pilot budget?", "identity": "demo-user", "maxTokens": 1200 }' ``` `/recall` returns ranked documents with source excerpts and no generated answer. A RAG pipeline can merge these documents with results from its existing retriever, then apply its own sufficiency policy. See the [API reference](/docs/memory-api/api-reference) and [benchmark method](/benchmarks). ## Frequently asked questions ### What is the best vector database for RAG? There is no single best option. Consider pgvector when you already run Postgres, Pinecone for a fully managed service, Qdrant for open-source self-hosting with payload filtering, Weaviate for combined keyword and vector search, and Milvus for distributed indexes. Compare deployment, filtering, scale, data location, and procurement requirements. ### Is pgvector good enough for production RAG? For most workloads, yes. It supports exact and approximate search with HNSW and IVFFlat indexes, and keeps vectors in the same database, transaction and backup as your rows. Move to a dedicated engine when you can name the constraint you hit: index size beyond one node, a filtering pattern that collapses recall, or a need to scale search independently of the rest of the database. ### Do I need a vector database for RAG at all? Not always. Under a few thousand documents, keyword search or a SQL filter may be sufficient and easier to debug. Use vectors when queries differ from the source wording and the corpus is too large to scan. ### Vector database vs RAG: are they the same thing? No. RAG is the pattern of retrieving context at question time and putting it in the prompt. A vector database is one component that can serve the retrieval step, alongside keyword search, SQL and API calls. ### Can a memory API replace a vector database? Only for the part of the corpus that is timestamped history. Documents you own, product content and anything static still belong in your own index. past.dev is built for the history where the answer changed, and `/recall` returns evidence in a form you can merge with the retriever you already have. ## Related - [Vector database vs graph database vs memory](https://past.dev/vector-database-vs-memory) - [What is agent memory?](https://past.dev/what-is-agent-memory) - [Context engineering for AI agents](https://past.dev/context-engineering) - [Memory API quickstart](https://past.dev/docs/memory-api/quickstart) - [How we measure memory](https://past.dev/benchmarks)