Integrating vector search capabilities into PostgreSQL via pgvector has democratized the deployment of Retrieval-Augmented Generation (RAG) systems and similarity search applications. However, moving from a proof-of-concept prototype to a high-throughput production environment often reveals performance bottlenecks. A naive implementation can lead to slow queries, high latency, and unnecessary CPU load. This post explores how to optimize pgvector performance through intelligent indexing strategies and query tuning.
Understanding the Indexing Landscape
pgvector currently supports two primary indexing methods: HNSW (Hierarchical Navigable Small World) and IVFFlat (Inverted File with Flat Quantization). Choosing the right index is the single most impactful decision you can make for performance.
HNSW is generally recommended for most production use cases. It offers superior recall rates and faster query times compared to IVFFlat, especially as dataset size grows. While it consumes more memory and takes longer to build, its ability to navigate the graph structure efficiently makes it ideal for low-latency requirements. Conversely, IVFFlat is lighter on memory and builds faster but requires more disk I/O and generally offers lower recall unless the number of lists is increased significantly. It is best suited for scenarios where memory is constrained or datasets are static and moderate in size.
Creating Optimized Indices
When creating an index, the choice of distance metric and specific parameters dictates the quality of the search. For text embeddings derived from models like BERT or OpenAI's text-embedding-3-small, cosine distance is often the standard. Here is how to create an optimized HNSW index:
CREATE INDEX CONCURRENTLY embedding_idx
ON documents
USING hnsw (embedding vector_cosine_ops);
However, default settings are rarely optimal for production. The m parameter controls the number of bi-directional links created during graph construction, while ef_construction influences the search quality during index building. For high-quality results, consider increasing ef_construction:
CREATE INDEX CONCURRENTLY embedding_idx
ON documents
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
Higher values for m and ef_construction increase memory usage and build time but significantly improve recall accuracy.
Tuning Query Performance
Even with a well-optimized index, query performance can degrade if the search parameters are not tuned. The hnsw_ef_search setting acts as a budget for the search algorithm. By default, it is set to 40, which balances speed and accuracy. For production workloads, you should tune this based on your acceptable latency and desired recall.
If you are experiencing high p99 latency, try reducing hnsw_ef_search. If your application requires near-perfect retrieval accuracy, increase it. You can set this at the session level:
SET hnsw.ef_search = 128;
SELECT *
FROM documents
ORDER BY embedding <#> '[0.1, 0.2, ..., 0.1]'::vector
LIMIT 10;
Practical Considerations for Scale
As your dataset grows beyond hundreds of millions of vectors, consider partitioning your tables by a non-vector column, such as a tenant ID or date. This allows you to create smaller, more focused indexes, reducing the overhead of graph traversal. Additionally, ensure that your database server has sufficient shared memory (shared_buffers) and WAL level settings configured to handle the heavy write load associated with index builds.
Conclusion
Optimizing pgvector is not a one-time setup but an iterative process. Start with HNSW for its balanced performance profile, tune ef_construction and ef_search to match your latency and accuracy SLAs, and monitor the trade-offs between memory usage and query speed. By applying these strategies, you can leverage the power of vector search within PostgreSQL without compromising on performance or reliability.