Vector Database for RAG: 5 Benchmarks to Run First
TL;DR
There is no best vector database for RAG. There is the one that fits your questions, your data and your budget.
RAG stands for retrieval augmented generation. It is a way to make an AI model answer from your own material, not from its general training. The vector database is the part that finds the right material.
If you run a coaching, education or consulting business, this choice decides whether your AI answers from your frameworks or from guesswork. You do not need to be an engineer to ask for the right tests. You do need to know what they are before you sign off the build.
Test five things on your own data first: retrieval quality, speed, filters, embedding time and cost. This guide covers each test, what the research measured, and when you do not need a dedicated store at all.
What a vector database does in a RAG system
A vector database is the memory layer. It stores embeddings, which are lists of numbers that capture the meaning of a piece of text. When the model needs context, it hands back the closest matches.
Think of a filing cabinet sorted by meaning, not by name. You ask a question. It pulls the files most likely to help, even when they share no words with what you asked.
The flow is short:
- Someone asks a question.
- A model turns the question into an embedding.
- The database finds the stored embeddings closest to it.
- A reranker, if you use one, puts the best results first.
- The top results go into the prompt with the question.
Step five is why this matters. A large language model in a business is only as good as the material it is handed. Give it your real frameworks and it answers like you. Give it nothing and it guesses.
Devwiz has a plain primer on what RAG is if you want the longer version.
There are three ways to search.
| Search type | What it finds | Where it fits |
|---|---|---|
| Vector search | Results with a similar meaning | Loose questions where the exact words do not matter |
| Keyword search | Exact matches | Product codes, names and legal terms |
| Hybrid search | Both, merged into one list | Most real business questions |
Hybrid tends to be the safer default. Take a client asking about the "refund policy for annual plans". The system needs the idea, which is refunds. It also needs the exact phrase, which is annual plans. Search by meaning alone can miss the second part.
The same embedding idea now shapes how AI search picks which pages to cite. BabyLoveGrowth covers that side in its guide to vector search SEO.
Five tests to run before you choose
Picking from a leaderboard screenshot is how teams end up rebuilding the retrieval layer six months later. Run these five tests first.
- Retrieval quality. Measure recall@k. That is how often the right document shows up in your top k results. Then check how much a reranker lifts it.
- Speed under load. Measure p95 latency, the time that 95 in 100 requests finish within. Do it with many users at once, not one clean query.
- Filters. Test searches with a condition attached, such as "similar documents, but only from last quarter". Each database handles these differently.
- Embedding time. Time how long it takes to turn a question into an embedding. This step is often slower than the search.
- Cost. Add up reads, writes, storage and index rebuilds. Pinecone's pricing page shows the shape: storage, write units and read units are each billed as their own line.
What the research measured
The best independent evidence here is BigVectorBench, a 2025 paper in the Proceedings of the VLDB Endowment. The authors tested Milvus, Weaviate and Qdrant.
Two findings stand out.
Compound queries hurt. A compound query is a search with extra conditions or more than one kind of input. In the paper's Milvus test, compound queries cut recall by 48.54% and throughput by 23.38%, against simple single-vector queries.
Embedding is the slow part. The authors found the main bottleneck in serving a query is embedding time, not the search. Embedding a document of 8,192 tokens took as long as 400 milliseconds on a GPU. The search in Milvus took about ten.
Filters are less predictable than you would guess. On one workload in the paper, adding a filter lifted the best recall from 0.91 to 0.95. On another database, filter queries held recall steady while throughput fell to almost zero. The same feature produced opposite results.
That is the whole case for testing on your own data. A database that looks great on ANN-Benchmarks, the standard public comparison, can behave very differently once your filters and your mix of content are in play. The paper argues for testing embedding and compound queries together, in one benchmark.
How to test a vector database for RAG on your own data
Public leaderboards test clean, simple queries. Your system will not look like that. Build a small test of your own.
Make it look like a normal day:
- Questions with conditions attached, not only plain ones.
- Writes and reads at the same time.
- Images, tables or code, if your material includes them.
- Several users at once.
- New content arriving while people search, to check for stale answers.
Then time the whole request, with the embedding step inside it. Record recall on your own questions, with answers you marked by hand. Someone else's dataset tells you about their problem.
A workable order:
- Start with a small sample of 1,000 to 10,000 documents. Obvious problems show up fast.
- Scale to the volume you expect and run the same tests again.
- Add users until you reach your expected peak.
- Run it with a reranker and without one. Compare the gain in quality with the extra wait.
When you read the results, look for trade-offs. Do not look for a winner. The fastest system at low volume can get slow or costly at ten times the data. Work out your cost curve from your own read and write patterns, not from a vendor's example.
The BigVectorBench code is public on GitHub if your team wants a starting point.
Who runs it: managed, self-hosted or unified
Your choice here decides who does the work. You, or the vendor.
| Option | What you get | What you give up |
|---|---|---|
| Managed service | Scaling, patching and uptime handled for you | Some control, plus fees for requests and data leaving the system |
| Self-hosted | Full control over data, settings and hardware | Your team owns backups, scaling, monitoring and upgrades |
| Unified platform | Cache, vector store and database in one system | Your scaling and your failures are tied to one vendor |
Redis lays out the trade-offs in its comparison of open source vector databases. Read it knowing that Redis sells the unified option.
Before launch, answer five questions:
- How often do you rebuild the index, and how long does it take?
- What is the restore plan if the store fails?
- Which numbers do you watch, and what triggers an alert?
- Does it scale on its own, or do you plan for peak by hand?
- Who can read the raw material behind the embeddings, and is that logged?
Storage is the other lever. Amazon S3 Vectors keeps vectors in object storage. AWS says it cuts the cost of storing and querying vectors by up to 90%, with sub-second queries. AWS pitches it for large archives that are not queried often. Keep the content people ask about every day in a fast store. Push old content to the cheap one.
How the pieces fit in production
Most production RAG systems share one shape. You chunk the documents, embed the chunks and index them. At question time you retrieve, rerank and pass the results to the model.
Any step can be the slow one. When something drags, time each step on its own.
Three patterns come up again and again.
Hybrid search. A vector index handles meaning. A full-text index handles exact words. The system merges the two lists. VeloDB's guide to RAG with hybrid search shows both searches running in a single SQL query, with the rankings fused.
Tiered storage. Hot vectors sit in memory for speed. Cold vectors sit in object storage for cost. AWS has a guide to choosing an AWS vector database for RAG that compares its own options.
Embed ahead of time. Embedding your documents once, up front, means the only embedding at question time is the question itself. Re-embed when the source changes.
How you chunk matters as much as where you store. A 40-page deck saved as one file is close to useless to an agent. Our guide to building a consultant knowledge base covers chunking and tagging. Devwiz walks through the full build in how to build a RAG pipeline.
A selection plan for a small team
- Write down real questions. Use the ones clients and staff ask now, not ones you imagine. Set a target for speed and a target for recall.
- Build a small test set. Use your own material. Mark the right answers by hand. Include questions with conditions from the start.
- Test two or three candidates. Same data, same questions, side by side.
- Pilot with real traffic. Start with a small slice. Watch the gap between the test numbers and what happens live.
- Plan rebuilds and backups before launch. Not after something breaks.
Give the pilot at least a week. A test run misses the slow drift that real use shows.
The part most teams get wrong
Most teams compare index algorithms. Few test the questions their users will ask.
Start with the questions that carry the most risk. Refunds. Compliance. Anything a client acts on. Those paths deserve more checking than a place on a leaderboard.
Many founder-led businesses do not need a dedicated vector database on day one. We said the same about AI agent memory. Start small. A single-file store or a Postgres extension such as pgvector will carry a small build. Move up when you serve many users at once.
In our own stack, the memory layer is a wiki. It is a living map of how the business connects. Agents read that map before they act. The wiki is indexed in Pinecone, a vector database, so the AI can find the right knowledge by meaning and not only by keyword.
The database was never the hard part. The map was. Structure beats raw intelligence.
That is the point of an AI Operating System. Your IP sits in the memory layer. AI employees built with Claude Code read it before they act. So the work scales without you in every loop. Njin has a plain account of what an AI operating system for business is. It is also what we build with founders in the 90-day program.
So choose the store last. Decide what the system has to know first. Then run the five tests on the two or three options that fit.
James Killick
Find out what your system needs to know
A vector database is only worth choosing once you know which decisions still run through you. Start there.
Take the Founder Bottleneck Assessment. It scores five dimensions in six minutes, then names your next move.
Sources
Frequently Asked Questions
James Killick
Founder
The AI Orchestrator. 10+ years building digital products and 200+ apps shipped, now helping $1M+ educators and consultants turn their IP into AI-powered delivery systems.
James Killick founded and runs The AI Orchestrators.
More from James Killick