A vector database is a system that stores content as embeddings, lists of numbers produced by an AI model to represent meaning, and retrieves the items whose meaning is closest to a query. Where a traditional search looks for matching words, a vector search can find a passage about “ending the contract early” when someone asks about “termination for convenience.” Vector databases are a common building block in retrieval-augmented generation (RAG), where an AI assistant looks up relevant company content before answering.
At a glance
- A vector database stores embeddings, numeric representations of the meaning of text, images or other content.
- It finds the items most similar to a query, which supports search by meaning instead of exact wording.
- Its most common business use is supplying relevant context to a large language model in RAG systems.
- Vector search is offered both by dedicated vector databases and as a feature of many existing databases, search engines and AI platforms.
- The stored data is derived from your content, so access control and data governance apply to it as much as to the source.
What problem it solves
Large language models (LLMs) on their own know what was in their training data, which doesn’t include your contracts, policies, tickets or product manuals. Pasting whole document libraries into every request isn’t practical. Something has to find the few passages that matter for each question.
Keyword search can do part of that job, but it struggles when people phrase things differently from the documents. Vector search finds content by meaning, so a question and a relevant passage can match even with no words in common. A vector database makes this fast across millions of items, which is what lets an internal AI assistant answer from your own content in a few seconds. The same technique supports recommendations, finding duplicate records and matching similar support cases.
How it works
Chunking. Source documents are split into smaller pieces, such as paragraphs or sections, so that search can return just the relevant part.
Embedding. Each piece is passed through an embedding model, which turns it into a vector, a list of hundreds or thousands of numbers. Pieces with similar meaning end up with similar vectors.
Indexing and storage. The database stores each vector with its original text and metadata, such as source, date, department and access permissions, and builds an index that makes similarity search fast. Many indexes use approximate methods that trade a small amount of accuracy for large gains in speed.
Querying. When a user asks a question, it is turned into a vector with the same embedding model, and the database returns the closest matches. Many systems combine this with keyword search and metadata filters, for example only documents this user may see, or only current policies.
Use by AI. In a RAG system, the retrieved passages are added to the prompt so the generative AI model answers using them and can show which passages it used, so people can check the evidence.
When it matters for buyers
- When building an internal AI assistant or search. The quality of retrieval often decides whether answers are useful, more than the choice of language model.
- When choosing a platform. Decide whether vector search should live in a dedicated database, your existing database or search engine, or a managed AI platform’s built-in store. Providers also use “vector store” for overlapping components, from full databases to lightweight indexes or software libraries, whose persistence, filtering, access control and scale vary, so check what a given product actually includes.
- When data is sensitive. Embeddings and stored text are copies of your content; where they are hosted matters for private AI and data residency.
- When permissions vary. Search must respect who can see which documents, or the AI may reveal content to the wrong people.
- When content changes often. Plan how new, edited and deleted documents are re-embedded, as part of data governance.
For help planning AI search and data platforms, see our artificial intelligence overview.
Questions to ask vendors
- Is vector search a dedicated product or a feature of a database or platform we already use?
- Where are embeddings and source text stored, in which region, and are they encrypted?
- How does search enforce our existing document permissions for each user?
- Which embedding models are supported, and what happens if we change models later?
- How are updates and deletions in source systems reflected, and how quickly?
- Does it support combined keyword and vector (hybrid) search and metadata filters?
- How is it priced: by data stored, number of vectors, queries or compute?
- Can we export our data and indexes if we move to another platform?
How it differs from retrieval-augmented generation (RAG)
Retrieval-augmented generation (RAG) is an overall approach: look up relevant content, then have a language model answer using it. A vector database is one component that can do the looking up. RAG systems can also retrieve with keyword search, a regular database query or an API call, and many use a mix. A vector database also differs from a traditional relational database, which is built for exact matches on structured fields such as an account number or date; vector databases are built for similarity of meaning, and many modern systems combine both.
