Once you have embeddings, you need to find the ones most similar to a query. That's nearest-neighbour search, and vector databases make it fast at scale.
Exact Versus Approximate Search
- Exact search compares the query with every vector. Perfectly accurate and fine for tens of thousands of vectors with NumPy.
- Approximate nearest neighbour (ANN) search uses clever indexes to find very close matches much faster, trading a little accuracy for speed at millions of vectors.
Common Index Types
- HNSW (hierarchical navigable small world graphs) — fast and accurate, widely used.
- IVF (inverted file) — clusters vectors and searches only the nearest clusters.
- Product quantisation — compresses vectors to save memory.
The Options
- Libraries: FAISS, Annoy, hnswlib — you manage storage yourself.
- Database extensions: pgvector for PostgreSQL, vector search in Elasticsearch and OpenSearch — keep vectors alongside existing data.
- Dedicated vector databases: managed services and open-source systems built for vector workloads.
Features to Look For
Metadata filtering (by user, date, permission), hybrid keyword + vector search, updates and deletes, and backup and access control.
Do You Need One?
For a prototype or a few thousand documents, an in-memory array is simpler. If you already run PostgreSQL, an extension may be enough. Choose a dedicated system when scale, filtering or operational needs demand it.
Evaluate Retrieval Quality
Whatever you use, test with realistic queries: does the right item appear in the top results?