AI Solutions

Vector Databases Compared: pgvector, Qdrant, Weaviate, Milvus & More

Choosing a vector database for RAG: pgvector, Qdrant, Weaviate, Milvus, Pinecone and Chroma compared on hosting, filtering, hybrid search, scale and cost model.

GPTLabAI team 7 min read

Choosing a vector database comes down to five questions: where it runs (self-hosted or managed), how well it combines vector search with metadata filters, whether it does hybrid keyword-plus-vector search, how far it scales, and how you pay for it. For most business RAG systems we start with pgvector in an existing PostgreSQL database. We move to Qdrant, Weaviate or Milvus when scale or search features require it, and to Pinecone when a team wants fully managed infrastructure. Here is how the options compare as of September 2026.

What a vector database does in a RAG system

In retrieval-augmented generation, documents are split into chunks, turned into embedding vectors and stored. At question time, the query is embedded and the database returns the most similar chunks, usually through an approximate nearest neighbour (ANN) index such as HNSW. Then an LLM writes the answer.

The vector index is only part of the job. In real systems we also need:

  • Metadata filtering: only search documents this user may see, for this customer, in this language.
  • Hybrid search: combine semantic similarity with keyword matching (BM25 or sparse vectors). Pure vector search often misses exact product codes, names and error messages.
  • Updates and deletes: documents change, and deleted data must actually disappear (important under GDPR).
  • Operational basics: backups, access control, monitoring.

If you are new to RAG, read our RAG best practices first. Chunking and retrieval strategy usually matter more than the database brand.

The main options at a glance

Database Licence Hosting Hybrid search Filtering Best fit
pgvector Open source (PostgreSQL extension) Self-host or any managed Postgres that ships it Vector + Postgres full-text search, combined in SQL Full SQL WHERE, joins, row-level security Teams already on Postgres; small to mid-size corpora
Qdrant Apache 2.0 (Rust) Self-host (Docker) or Qdrant Cloud Dense + sparse vectors Rich JSON payload filters Performance-focused self-hosting, strong filtering
Weaviate BSD 3-Clause (Go) Self-host or Weaviate Cloud BM25 + vector built in Structured filters, multi-tenancy Built-in vectoriser modules, multi-tenant SaaS
Milvus Apache 2.0 (LF AI & Data) Lite / Standalone / Distributed, or Zilliz Cloud Sparse vectors + BM25 full-text Metadata filtering, multi-tenancy Very large scale, Kubernetes-native, GPU indexing
Pinecone Proprietary Fully managed (serverless); BYOC option Dense + sparse / full-text in one index Metadata filters Teams that want zero database operations
Chroma Apache 2.0 In-memory, local, client-server, or Chroma Cloud Full-text + vector Metadata and document filters Prototypes, local tools, simple apps

Other options you will come across include LanceDB (embedded, file-based), Elasticsearch/OpenSearch (if you already run them for search), Redis, and vector search built into MongoDB Atlas and the major cloud databases. If your team already operates one of these well, its vector features may be enough.

pgvector: the sensible default

pgvector adds a vector column type and ANN indexes (HNSW and IVFFlat) to PostgreSQL. It also supports halfvec (half precision), sparsevec and binary vectors, and several distance metrics. Recent versions add iterative index scans, which fix the old problem of filtered queries returning too few results.

Why we reach for it first:

  • One database. Documents, metadata, permissions and vectors live together. Filtering by tenant or access rights is a normal SQL WHERE clause or row-level security policy.
  • Transactions and deletes work the way your team already expects.
  • Hybrid search is possible by combining pgvector with Postgres full-text search, for example with reciprocal rank fusion in SQL.
  • Easy to host. It runs wherever Postgres runs, including your own servers. That matters for private RAG.

Limits to know: indexed vector columns support up to 2,000 dimensions (4,000 with halfvec), and large HNSW indexes need a lot of memory and take time to build. At tens of millions of vectors with heavy query load, a dedicated engine usually becomes easier to run.

Qdrant: fast, filter-friendly, self-hostable

Qdrant is a dedicated vector engine written in Rust. Strong points:

  • Payload filtering designed to work with the vector index rather than as a post-filter.
  • Dense, sparse and multi-vector support for hybrid and late-interaction retrieval.
  • Quantization to cut memory use.
  • Distributed mode with sharding and replication.
  • A single Docker container to start, and a managed cloud with a free tier.

It is a good choice when you want a dedicated vector store you can self-host, with heavy filtering (multi-tenant apps, permission-aware search).

Weaviate: batteries included

Weaviate stores objects and vectors together and ships hybrid search (BM25 + vector), built-in vectoriser modules that call embedding providers for you, reranking and multi-tenancy. It offers REST, gRPC and GraphQL APIs. It fits teams that want the database to handle embedding and hybrid ranking with little glue code, or SaaS products that need clean per-tenant isolation.

Milvus: built for very large scale

Milvus is a distributed, Kubernetes-native system with a wide choice of index types (HNSW, IVF, DiskANN and others), GPU acceleration, hot/cold storage and native BM25 full-text search through sparse vectors. Milvus Lite runs inside a Python process for development, and Zilliz Cloud is the managed version. Choose it when you expect hundreds of millions to billions of vectors and have the platform skills to operate it, or pay for the managed service.

Pinecone: fully managed

Pinecone is a proprietary, managed-only vector database with a serverless architecture. You do not run servers. Its current platform supports metadata filtering, dense and sparse vectors and full-text search in one index, plus integrated embedding and reranking. There is also a bring-your-own-cloud option for running inside your own cloud account. It suits teams without database operations capacity who are comfortable with a SaaS vendor holding their vectors (check the available regions) and a usage-based bill.

Chroma: easiest to start

Chroma is the fastest way to get vectors into a prototype: pip install chromadb and you have an in-memory or local persistent store. It now also offers a client-server mode and a hosted cloud with full-text, regex and metadata filtering. It is great for notebooks, internal tools and small apps. For large multi-tenant production systems, compare it carefully with the options above.

Understanding the cost model

We will not quote prices here, because they change often. Check each vendor’s current pricing page. What matters is how you pay:

Model Who You pay for Watch out for
Self-hosted open source pgvector, Qdrant, Weaviate, Milvus, Chroma Servers (RAM matters most for HNSW), storage, your team’s ops time Memory sizing, backups, upgrades
Managed cluster Qdrant Cloud, Weaviate Cloud, Zilliz Cloud, managed Postgres Provisioned capacity per hour/month Paying for idle capacity
Serverless / usage-based Pinecone, serverless tiers of others Storage per GB, read and write units, sometimes inference tokens Bills that grow with query volume; plan minimums

Rules of thumb for estimating:

  • Memory is the main cost driver for HNSW indexes. It scales with vector count × dimensions × bytes per dimension, plus index overhead. Smaller embedding dimensions, half precision and quantization reduce it a lot.
  • Embedding generation is often a larger cost than storage. Re-embedding your whole corpus when you change models is a real expense, so plan for it.
  • Operations time is a real cost of self-hosting. It is small for pgvector on an existing database and larger for a distributed Milvus cluster.

For the LLM side of the bill, see our LLM cost optimization guide.

How to choose: a decision guide

  1. Already on PostgreSQL, under a few million chunks? Use pgvector. Revisit only if you hit clear limits.
  2. Need data on your own infrastructure with strong filtering and more scale? Qdrant or Weaviate, self-hosted.
  3. Hundreds of millions of vectors or more? Milvus (or Zilliz Cloud), or a managed service sized for it.
  4. No ops capacity, speed to market matters most? Pinecone or a managed cloud tier of an open-source engine.
  5. Prototype or local tool? Chroma, or pgvector in a Docker container.
  6. Strict EU data residency or on-premise requirements? Prefer a self-hostable open-source option, and check the region options of any managed service.

Whatever you choose, keep the vector store behind a small internal interface in your code, and keep the source documents and embedding pipeline reproducible. Switching databases later is then a migration project, not a rewrite.

Key takeaways

  • Retrieval quality depends more on chunking, hybrid search and filtering than on the database brand.
  • pgvector is the pragmatic default for most business RAG systems.
  • Qdrant and Weaviate are strong self-hostable engines. Milvus targets very large scale. Pinecone removes operations work. Chroma is ideal for getting started.
  • Compare cost models (self-hosted, provisioned or usage-based) rather than headline prices, and include embedding and operations costs.
  • Keep your data pipeline portable so you can change databases later.

If you are planning a RAG system and want help choosing and setting up the right retrieval stack, especially one that keeps data on your own infrastructure, see our private RAG service or contact us.

7 min

How Much Does an AI Chatbot for Business Cost?

What drives AI chatbot cost for a business: build, model API usage, hosting and maintenance, build vs buy, plus a simple formula to estimate your token costs.

Read article

Have a project in mind? Let’s talk.

Whether you run a business or a research group, tell us what you need built, fixed or evaluated. You get a free consultation and a clear written estimate — no obligation.

  • Free consultation
  • Written scope and estimate
  • We reply within one working day
Contact us