Is the Vector Database Dead?
Dedicated vector databases are not dead, but they are no longer the default answer. Many teams now store embeddings in databases they already run, combine vector and keyword search, or let agents search files directly. A dedicated vector database still earns its place at large scale, with strict latency needs, or with heavy filtering and multi-tenant isolation.
Why people say it is dead
Three developments drive the claim. First, mainstream databases added vector support, so teams can keep embeddings next to their other data without a new system. Second, context windows grew large enough that small document sets can sometimes be passed to a model directly. Third, coding agents showed that searching files with ordinary tools, such as grep, can work well for code, where exact identifiers matter more than semantic similarity.
Each point is real. None of them makes vector search obsolete; they change where a dedicated vector database is the right tool.
What is actually true
Vector search is everywhere. Semantic retrieval underpins most production RAG systems, recommendation, and deduplication. What has changed is where the vectors live.
Hybrid search wins often. Combining keyword search, which handles exact terms, names, and codes, with vector search, which handles meaning and paraphrase, usually retrieves better than either alone. Many teams get more from adding keyword search and reranking than from switching vector stores.
Long context has limits. Large contexts are expensive per request, slower, and models use information in very long contexts unevenly. For large or frequently changing corpora, retrieval remains cheaper and more accurate.
Agentic search complements retrieval. Agents that grep and read files work well on codebases and structured folders. They are slower and costlier on large, unstructured document collections, where an index answers in milliseconds.
When a dedicated vector database earns its place
- Scale. Hundreds of millions of vectors or more, where specialised indexing and memory management matter.
- Latency at volume. High query rates with tight latency targets.
- Heavy filtering. Complex metadata filters combined with similarity search, done efficiently.
- Multi-tenancy. Strict isolation between many customers' data.
- Specialised features. Advanced index types, quantisation options, or hybrid scoring built in.
When something simpler is enough
- A vector extension in a relational database you already operate, for moderate scale.
- A search engine with vector support, when keyword search is already central.
- In-memory or file-based indexes for small, mostly static corpora.
- Direct file search by agents for code and well-structured folders.
Fewer systems means fewer things to secure, back up, monitor, patch, and keep consistent.
A decision guide
Start from your corpus and queries. If you have a few thousand documents that rarely change, a vector extension in your existing database, or even an in-memory index, is enough. If most queries contain exact identifiers, product codes, or names, invest in keyword search first and add vectors for paraphrased questions. If you serve many customers with isolated data, strict latency targets, and hundreds of millions of vectors, a dedicated vector database is likely worth its operational cost. Revisit the choice when scale or query patterns change, not when a new trend appears.
Cost and operations
Every additional data store adds work: provisioning, backups, upgrades, security reviews, access control, and keeping its contents consistent with the source of truth. That cost is easy to underestimate when a vector database is chosen for a prototype and then becomes part of production. Keeping embeddings in a database you already run avoids most of it, until scale or features justify the extra system.
What matters more than the store
Retrieval quality depends mostly on chunking, embedding model choice, hybrid search, reranking, metadata, and keeping the index current. Teams that switch vector databases to fix poor retrieval usually find the problem moved with them. Measure retrieval directly: for a set of real questions, do the passages containing the answer appear in the results? Fix that first, then choose the store.
Frequently asked questions
- Do I still need a vector database for RAG?
- Not necessarily a dedicated one. Many RAG systems store embeddings in a relational database or search engine with vector support. A dedicated vector database makes sense at large scale, with strict latency needs, heavy filtering, or multi-tenant isolation. Retrieval quality depends more on chunking, embeddings, and hybrid search than on the store.
- Can long context windows replace vector search?
- For small, static document sets, sometimes. For large or frequently changing corpora, retrieval is usually cheaper, faster, and more accurate, because very long contexts are expensive per request and models use information within them unevenly. Many systems combine retrieval with moderately sized contexts.
- Is grep better than vector search for code?
- For finding exact identifiers, function names, and error strings, text search is often better. For questions about behaviour or intent, such as where authentication is handled, semantic search helps. Coding agents commonly combine both, and an index of the codebase saves time and tokens on large repositories.
- What is hybrid search?
- Hybrid search combines keyword search, which matches exact terms, with vector search, which matches meaning, and merges or reranks the results. It handles queries containing names, codes, and specific terms better than pure vector search, and paraphrased questions better than pure keyword search.
- What is the best embedding model for RAG?
- It depends on your content, languages, and constraints. Evaluate a few strong candidates on your own data using real questions, measuring whether the right passages are retrieved, and weigh dimension size, cost, latency, and whether you can self-host. Retrieval quality on your corpus matters more than public benchmark rankings.