A vector database stores data as high-dimensional numerical vectors and retrieves records by similarity rather than by exact value. Instead of asking "find the row where id = 42", you ask "find the rows most similar to this query". That makes it the standard storage layer for semantic search, recommendation and retrieval-augmented generation.
Why vectors, not rows
Most data that matters in AI pipelines, text, images, audio, has no natural primary key you can match against. What you want to know is: which stored items are close in meaning to this new input?
An embedding model converts each item into a vector: a list of numbers, often hundreds or thousands of dimensions long, where position in that space encodes semantic meaning. Two sentences that say the same thing in different words will produce vectors that are close together. Two sentences that mean opposite things will be far apart.
A vector database indexes those vectors so that nearest-neighbour queries run in milliseconds, even over millions or billions of records.
How similarity search works
The most common similarity measure is cosine similarity: the angle between two vectors. Others include dot product and Euclidean distance. Exact nearest-neighbour search over large datasets is computationally expensive, so most systems use approximate nearest-neighbour (ANN) algorithms. These trade a small, controllable amount of recall for large speed gains. Common indexing strategies include HNSW (Hierarchical Navigable Small World) and IVF (Inverted File Index).
You can usually combine vector similarity with conventional filters: "find the ten most similar documents, but only from this tenant's data, published after this date". That hybrid query is how most production systems use vector databases in practice.
Where vector databases appear in production
- RAG pipelines. Retrieved chunks of text, chosen by semantic similarity to the user's question, are passed to a language model as context. The vector database is what makes retrieval fast and relevant. See what is RAG for the fuller picture.
- Semantic search. Product catalogues, support knowledge bases, legal document stores: anywhere keyword search fails because users phrase things differently from how documents are written.
- Recommendation. User and item embeddings are stored; similarity search finds items close to what a user has engaged with.
- Deduplication and clustering. Finding near-duplicate records, or grouping similar items, without needing an exact string match.
Dedicated systems versus extensions
Some teams use purpose-built vector databases. Others reach for a vector extension bolted onto a relational database they already run. Both approaches are in production at scale. The right choice depends on query volume, latency requirements, whether you need strong transactional guarantees on the same records, and operational familiarity. There is no universal answer.
The context window connection
Vector search is partly a workaround for the limits of context windows. You cannot pass an entire knowledge base to a language model; you retrieve the relevant slices and pass those. As context windows grow, the calculus shifts, but retrieval remains faster and cheaper than stuffing everything into a prompt, and it keeps the source of truth outside the model.
What we test for
Engineers working on RAG or semantic search pipelines need to understand how to call a vector database API and, more importantly, how to evaluate whether retrieval is actually working: whether the embeddings chosen are appropriate, whether chunking strategy affects recall, and when a model-based approach is the wrong tool entirely. Our vetting covers evaluation design and retrieval quality judgement as part of the AI engineer assessment. Full detail at how we vet.
Short answers
What is the difference between a vector database and a regular database?
A regular database retrieves records by exact value or range. A vector database retrieves by similarity: which stored vectors are closest to a query vector. They answer different questions and are often used together in the same system.
Do you need a dedicated vector database, or can a relational database handle it?
Both work in production. Dedicated systems typically offer better performance at high query volume. Relational extensions are simpler operationally if you already run that database. The right choice depends on scale, latency needs and team familiarity.
Is a vector database the same as a RAG pipeline?
No. A vector database is one component inside a RAG pipeline. RAG also involves an embedding model, a chunking strategy, a language model and an orchestration layer. The database handles storage and retrieval; the rest of the pipeline assembles and uses what it returns.