AI · RAG · Vector Databases · Engineering

You Don't Need a Dedicated Vector Database. You Need to Know When You Will.

7 October 2026 · Devectra

The default answer nobody likes

Every RAG architecture conversation eventually turns into "which vector database should we use," and the honest answer for most teams is: the one you already have. If your application runs on Postgres, pgvector is a one-line extension install, not a new system to provision, back up, monitor, and staff for. It inherits your existing replication, your existing access control, your existing on-call runbook. A dedicated vector database adds all of that back from zero, for a problem you may not have yet.

That's not a popular answer in a market where a dozen vendors are selling the dedicated option, but it's the one borne out by how RAG systems actually fail: almost never because the index was the wrong shape, and almost always because retrieval quality was never measured, chunking was naive, or metadata filtering was an afterthought. Database choice is rarely the bottleneck. It becomes one at a specific, nameable scale — and knowing where that line sits is the actual skill.

Where pgvector holds and where it doesn't

With HNSW indexing, pgvector handles most production workloads cleanly under roughly 10 million vectors, with query latency in the 5–20ms range on reasonable hardware. That covers the overwhelming majority of SMB and institutional use cases — a document corpus, a knowledge base, a support archive, a course catalog. None of that gets anywhere near 10 million chunks.

Past that point, the picture changes in stages. Between roughly 10M and 50M vectors, keeping pgvector fast is a real tuning job: quantized vector types to cut memory pressure, careful index-build memory budgeting, and often a disk-backed ANN extension to avoid holding the whole index in RAM. Above 50M, running pgvector well is a genuine infrastructure investment — still frequently the right call on cost, but no longer a "just add an extension" decision.

Dedicated vector databases earn their keep at that scale: distributed architecture that shards cleanly across machines, real-time index updates without rebuild windows, and metadata filtering built into the index itself rather than bolted on as a post-filter. If you're genuinely approaching tens of millions of embeddings with write-heavy update patterns, that operational model is worth paying for — and dedicated vector services typically cost several times more than the incremental Postgres spend, which is the price of someone else owning that distributed system for you.

The failure mode we see more often than "we outgrew pgvector" is the opposite: a team adopts Pinecone or Qdrant on day one for a corpus of forty thousand documents, then spends engineering time keeping two databases — the vector store and the system of record — consistent with each other. That synchronization problem, not query latency, is what actually costs time.

The embedding model is doing more work than the database

Teams spend disproportionate energy on the database and comparatively little on the embedding model, which is backwards — the model determines what "similar" even means before the database does anything with it. OpenAI's text-embedding-3-small runs around $0.02 per million tokens at 1536 dimensions; text-embedding-3-large runs about $0.13 per million tokens at 3072 dimensions for meaningfully better retrieval accuracy on harder queries. Voyage AI's models post some of the strongest public benchmark results for retrieval specifically, at a comparable per-token price. The gap between a mediocre embedding model and a strong one shows up as silently wrong retrieval — the kind that doesn't throw an error, it just quietly returns the fourth-most-relevant chunk instead of the first.

Dimension size is the other lever everyone underuses. Newer embedding models support truncating output dimensions — using 512 or 1024 of a 3072-dimension vector instead of the full thing — with a small, measurable accuracy cost and a large storage and compute win. For most SMB-scale corpora, that tradeoff is free money: smaller vectors mean faster search and lower memory pressure, and the accuracy delta is rarely the thing limiting your retrieval quality anyway.

Where naive implementations actually break

The vector search itself is rarely the problem we get called in to fix. The recurring ones are:

  • Metadata filtering treated as an afterthought. "Find similar chunks, then filter by department/date/permission" is a different, slower, and often wrong operation than filtering first and searching within that set. If your data has access control (a school's staff-only policies vs. parent-facing content, a company's internal vs. client-facing docs), the filter has to be part of the query plan, not a step after it.
  • No re-embedding plan. Source documents change. Embeddings don't update themselves. Teams build the initial pipeline, ship it, and six months later the vector store is quietly stale against the document system of record — with no process that caught it, because nothing errors when an embedding is just out of date.
  • One embedding model, every content type. A support ticket, a contract clause, and a product spec don't chunk or embed the same way. Treating them identically is a common source of retrieval quality that degrades unevenly across content types, which is exactly the kind of failure that's hard to spot in a demo and easy to spot in a client complaint three weeks after launch.

The actual decision framework

Start with pgvector if you already run Postgres and your corpus is in the low millions of chunks or fewer — which describes nearly every SMB, school, and institutional deployment we see. Move to a dedicated vector database when you can name the specific thing forcing the move: a vector count you've actually measured approaching the tens-of-millions range, a write throughput requirement pgvector can't hit, or a filtering pattern that's measurably slow in production, not in theory.

What you should not do is pick the database first and the embedding strategy second. The database is replaceable with a migration script. A bad embedding model or a chunking scheme that ignores your content's actual structure is a redesign, not a migration — and it's the one that determines whether retrieval works at all.

Takeaway: for the vector database decision, default to what you already run and change it only when you can point to the specific limit you hit. Spend the energy you were about to spend comparing vector databases on your embedding model and your chunking strategy instead — that's where retrieval quality actually lives.

Have a system like this in mind?

Get a scoped plan ↗