Skip to main content

Choosing a Vector Store for Spring AI: pgvector vs Redis vs Qdrant vs Elasticsearch

The same 30,000 documents through Spring AI 2.0’s VectorStore API into pgvector, Redis, Qdrant and Elasticsearch, all on library defaults: ingest time, recall@10 against exact search, latency, metadata filtering and running cost, and the one setting that fixes each surprise.

You have text, and you want to find the pieces of it that mean something close to a question. That is the job of a vector store, and Spring AI gives you one interface, VectorStore, in front of a long list of them. The interface makes switching look free: change a dependency, change a few properties, and the same add() and similaritySearch() calls run against a different database. It is not free, and the part that is not free is quiet. Every store has defaults, the defaults differ, and the differences show up as answers that are slightly wrong rather than as errors. This article puts the same 30,000 documents into four stores — pgvector, Redis, Qdrant and Elasticsearch — through Spring AI 2.0’s VectorStore API, leaves every store on the library’s defaults, and measures how fast each one ingests, how many of the true nearest neighbours each one actually returns, how each behaves when you add a metadata filter, and what each costs to keep running. Then it shows, for each surprise, the single setting that fixes it. Depth is in expandable sections; read straight through or open only what you need.
Versions this was written and run against. Spring Boot 4.1.1, Spring AI 2.0.1 and Java 25. The stores were real processes on one 2-vCPU machine, no Docker: PostgreSQL 16.15 with pgvector 0.8.7 (built from its v0.8.7 tag), Redis Stack 7.4.0-v8 (Redis 7.4.7 with RediSearch 2.10.20), Qdrant 1.19.2 and Elasticsearch 9.5.5 (single node, security off). All the code is in the vector-stores module of asmhatre/spring-ai; every console block below is quoted from a file under its output/ directory, written by a test that asserts the same numbers. The data is synthetic (clustered random vectors, no embedding model), so the numbers compare the stores; they say nothing about how any embedding model performs, and recall on real embeddings will differ — most of all for Elasticsearch’s quantised default.

A vector store answers one question: which of these is closest to that?

A language model can turn a piece of text into a list of numbers, usually a few hundred of them, called an embedding. The useful property is that texts with similar meaning get lists that point in similar directions. “How do I reset my password” and “I forgot my login” end up close together; “quarterly revenue” ends up far away. A vector store keeps those lists, and when you hand it the embedding of a question it returns the stored items whose lists are nearest. That is retrieval-augmented generation in one sentence, and it is why a store sits behind almost every Spring AI application that talks about your own data. Finding the exact nearest items means comparing the question against every stored vector. That is fine for a thousand documents and too slow for ten million, so real stores build an index that finds almost the nearest ones much faster. The four stores here all use the same family of index, HNSW (a layered graph that you walk towards the answer). “Almost” is the important word, and the rest of the article is about how much “almost” each store gives you by default.
Documentstext + metadataEmbeddingModeltext → 384 numbersVectorStoreadd() / similaritySearch()pgvectorRedisQdrantElasticsearchSame documents, same vectors, same calls. What differs is everything behind the interface: the index, the defaults and the filter.
The diagram is the whole experiment. Every store receives identical vectors from one stand-in EmbeddingModel that looks them up by name (d12 is document 12), so no store can win by being fed better data. The dashed lines are the part Spring AI hides from you, and the part this article opens.
Going deeper: the dataset, and what “recall@10” means here
Recall@10 is the share of the ten truly nearest documents that a store actually returned. The truth is computed by brute force in Dataset.java: for each of 100 queries, a plain loop scores all 30,000 vectors and keeps the ten best. A store that returns eight of those ten has recall 0.8. The stores are all asked for topK=10 with cosine distance. The dataset is 30,000 unit-length vectors of 384 dimensions grouped into 40 clusters, with three metadata fields per document: category (eight values, tied to the cluster), year (2015–2025) and region. It is synthetic because there is no embedding API key in the build and a downloaded model would make the repository depend on a model host. The cost of that choice is real: quantisation-based indexes in particular can behave differently on real embeddings, and the article says so where it matters.

One interface, four stores — and one thing the interface does not do for you

With Spring Boot’s starters you normally never build a store yourself; the auto-configuration does. To put four stores in one program, the module builds each by hand with its builder, using the library’s defaults wherever the builder lets it (the file is StoreFactory.java). Here is the pgvector one — the other three are the same shape.
    public static VectorStore pgvector(EmbeddingModel em) {
        JdbcTemplate jdbc = pgJdbc();
        jdbc.execute("CREATE EXTENSION IF NOT EXISTS vector");
        return init(PgVectorStore.builder(jdbc, em)
                .vectorTableName(NAME)
                .dimensions(Dataset.DIMS)
                .distanceType(PgVectorStore.PgDistanceType.COSINE_DISTANCE)
                .indexType(PgVectorStore.PgIndexType.HNSW)
                .removeExistingVectorStoreTable(true)
                .initializeSchema(true)
                .build());
    }
Building by hand exposed the first trap, which would also hit you in a unit test that constructs a store directly. The first run failed on the very first insert:
The table does not exist yet. build() returns a store but creates nothing. The pgvector table, the Redis index, the Qdrant collection and the Elasticsearch index are all created in afterPropertiesSet(), which Spring calls for a bean and nobody calls for a builder. My first run died with ERROR: relation "public.vs_bench" does not exist. The fix is the init() helper in StoreFactory.java, shown next, which calls it explicitly. In a normal Boot application you will never see this.
    static VectorStore init(VectorStore store) {
        try {
            ((org.springframework.beans.factory.InitializingBean) store).afterPropertiesSet();
        } catch (Exception e) {
            throw new IllegalStateException(e);
        }
        return store;
    }
If you are coming from Spring AI 1.x, two more things changed on the way in, both in the Redis store. I compiled the 1.x way of wiring it against 2.0.1 and kept the compiler’s own words:
$ javac -cp <spring-ai 2.0.1 classpath> src/broken/Redis1xStyle.java
src/broken/Redis1xStyle.java:11: error: cannot find symbol
                .vectorAlgorithm(RedisVectorStore.Algorithm.HSNW)
                                                           ^
  symbol:   variable HSNW
  location: class Algorithm
src/broken/Redis1xStyle.java:10: error: incompatible types: JedisPooled cannot be converted to RedisClient
        return RedisVectorStore.builder(jedis, em)
The builder now takes Jedis 7’s RedisClient rather than JedisPooled, and the enum constant that was spelled HSNW in 1.x is now HNSW. The file that produces this message is Redis1xStyle.java; it is deliberately outside the build.
Going deeper: why not just use the four starters?
Each store has a starter (spring-ai-starter-vector-store-pgvector and so on). Put all four on one classpath and Boot would try to create four VectorStore beans in one context. The module depends on the store artifacts (spring-ai-pgvector-store, spring-ai-redis-store, spring-ai-qdrant-store, spring-ai-elasticsearch-store) directly and builds each one itself, so the comparison is in one process. In your application you want exactly one starter. The versions come from the Spring AI BOM: Jedis 7.4.1, the Qdrant Java client 1.18.0 and the Elasticsearch Java client 9.4.5 (note the client is one minor version behind the 9.5.5 server I ran; it worked).

Getting the data in: the fastest and the slowest differ by more than an order of magnitude

The first measurement is simply how long it takes to add 30,000 documents in batches of 1,000. Embeddings come from the lookup, so this is time spent in the store and its index, not in an embedding API. Here is the transcript from 01-ingest.txt:
dataset: 30000 documents x 384 dims, 40 clusters, batches of 1000; embeddings are looked up, so this is store time only
store             seconds  docs/second
pgvector             46.5          645
redis                 9.5         3158
qdrant                1.7        17382
elasticsearch         9.9         3039
Qdrant took 1.7 seconds and pgvector 46.5, a ratio of about 27 to one. Redis and Elasticsearch landed together at about 9.5 seconds. The likely reason pgvector is slowest is that its HNSW index is maintained inside each insert, row by row, from the first row on — I did not isolate that. Spring AI’s store creates the index up front, so you cannot try the usual alternative (load first, build the index after) through the store; you would do that in plain SQL.
“Ingest finished” is not “searchable”. Elasticsearch needs a refresh before new documents appear, Qdrant runs an optimizer pass in the background, and Redis drains an indexing queue. A benchmark, or an integration test, that searches the instant add() returns can see a partial index. The test waits for each one (a _refresh, a green collection status, percent_indexed of 1) before it measures anything.
Going deeper: ingest in production
These are single-writer, single-node numbers on a shared 2-vCPU box, so read them as ratios, not as capacity. Two things change the picture in production. Embedding time usually dwarfs store time (a hosted embedding API at a few hundred documents per second is slower than any of these stores), so ingest speed rarely decides the choice. If pgvector’s per-row index maintenance is what makes it slow (unverified here), loading in bulk and indexing afterwards in plain SQL is the usual remedy for a large migration.

Finding the nearest ten: two of four stores miss half of them by default

Now the question that matters. 100 queries, ten results each, against the exact answer. Nothing is tuned: this is what you get by following each store’s reference page. The transcript is 02-unfiltered.txt:
100 queries after 20 warm-up, topK=10, cosine, each store's library defaults
store           recall@10    p50 ms    p95 ms  mean hits
pgvector            0.951       9.5      16.4       10.0
redis               0.616       0.7       1.1       10.0
qdrant              0.999       2.8       6.8       10.0
elasticsearch       0.491       5.4      20.5       10.0
pgvector0.95Redis0.62Qdrant1.00Elasticsearch0.49recall@10 on library defaults, 30,000 documents, 100 queries (higher is better)
Read the chart as a count of how many of the ten correct answers you get. Qdrant and pgvector return nearly all of them. Redis returns about six, and Elasticsearch returns about five — and neither tells you. Every query returned ten results; they were just not the ten closest. Latency tells a different story: Redis answered in under a millisecond (0.7 ms median), Qdrant in 2.8, pgvector in 9.5, Elasticsearch in 5.4. A fast wrong answer looks identical to a fast right one until you measure against the truth.

Redis: one runtime setting, and recall goes from 0.62 to 0.96

Redis’s HNSW index has a query-time knob, EF_RUNTIME, which is how many candidates the graph walk keeps before it stops. Left unset, Redis uses 10 — the same as the topK of 10 I asked for, which leaves no slack for the walk to recover from a wrong turn. Spring AI exposes it as hnswEfRuntime(...) on the builder. I rebuilt the index with three settings and loaded the same documents each time (06-redis-ef-runtime.txt):
Redis HNSW EF_RUNTIME: the library default (not set by Spring AI) against explicit values; index rebuilt and reloaded each time
ef_runtime    recall@10    p50 ms  filt recall    filt p50 ms
default           0.616       0.6        1.000            1.6
50                0.962       0.6        1.000            1.3
200               1.000       0.8        1.000            1.3
At 50, recall is 0.962 for no measurable cost in latency; at 200 it is 1.000. The fix is one line, .hnswEfRuntime(100), and I would set it explicitly in any Redis-backed application rather than inherit 10.

Elasticsearch: the default mapping is quantised, and on this data it costs half the recall

Elasticsearch’s result was the surprise, and the first explanation I wrote down was wrong, so here is the path. The Spring AI store asks for k=10 with num_candidates set to 1.5 times topK, which is 15 — I read that straight from the store’s bytecode (ldc2_w 1.5d, dmul, d2i). Fifteen candidates for ten results looks stingy, so I assumed that was the cause and tested it by sending the same query directly to Elasticsearch with num_candidates of 15, 50, 100 and 500 (05-elasticsearch-num-candidates.txt):
same index, same 100 queries, raw _search with the num_candidates varied:
num_candidates    recall@10    p50 ms
15                    0.491       4.5
50                    0.489       3.9
100                   0.490       3.8
500                   0.490       4.0
Recall did not move. Whatever limits it, it is not the candidate count. The next suspect was the index itself. The mapping the store creates, captured from the live index in 04-defaults-that-matter.txt, is:
elasticsearch embedding mapping: "embedding":{"type":"dense_vector","dims":384,"index":true,"similarity":"cosine","index_options":{"type":"bbq_hnsw","m":16,"ef_construction":100,"rescore_vector":{"oversample":3.0}}}
bbq_hnsw is what Elasticsearch applied here: its documentation makes it the default for vectors of 384 dimensions or more. It compresses each vector to about one bit per dimension (“better binary quantization”), searches the small copies, then re-scores the top candidates with the full vectors (oversample: 3.0). It saves a great deal of memory. To see what it costs here, I created two more indexes by hand that differ from the default in exactly one field, index_options.type, loaded the same documents through the same Spring AI store, and measured (10-elasticsearch-index-options.txt):
type          recall@10    p50 ms  filt recall    filt p50 ms      store
bbq_hnsw*         0.491       3.6        0.544            3.0     54.4mb
int8_hnsw         0.952       3.0        0.974            3.1     64.1mb
hnsw              0.995       3.0        1.000            2.7     52.6mb
* default; its full mapping is in 04-defaults-that-matter.txt
Same data, same store class, same queries: bbq_hnsw gives 0.491, 8-bit quantisation (int8_hnsw) gives 0.952, and unquantised hnsw gives 0.995. The index type is the cause. Two honesty points belong here. Elasticsearch builds its graph non-deterministically, so the default row moved between 0.47 and 0.50 over the runs I did; the gap to the others is far larger than that wobble. And this data is synthetic, a stress test for one-bit quantisation, because my vectors are random points around cluster centres with no fine structure to preserve. Elastic’s claim that BBQ holds recall on real embeddings is not something this experiment can confirm or refute. What it does show is that the default is a decision made for you.
What to do. Spring AI’s ElasticsearchVectorStoreOptions has no field for index_options, so you cannot choose the type through the store. Create the index yourself with the mapping you want (the test shows the PUT) and build the store with initializeSchema(false). Then measure recall on a sample of your embeddings before you trust either choice.
Going deeper: HNSW in two minutes, and the three numbers that matter
HNSW stores each vector as a node in a graph with several layers. A search enters at the sparse top layer, hops to the neighbour closest to the query, drops down a layer, and repeats until the dense bottom layer, where it keeps exploring until it has looked at ef candidates. Three numbers control the trade between speed and accuracy: m (links per node, fixed at build time), ef_construction (candidates considered while building) and ef at query time (called ef_search in pgvector, EF_RUNTIME in Redis, ef/hnsw_ef in Qdrant, and num_candidates in Elasticsearch). A larger query-time ef costs more time and finds more of the true neighbours. The defaults in this run were: pgvector m=16, ef_construction=64, ef_search=40; Redis EF_RUNTIME=10 (my reading of the library behaviour, confirmed by the recall change above); Qdrant m=16, ef_construct=100 (read from the live collection); Elasticsearch m=16, ef_construction=100.

Adding a metadata filter: the same query, four different behaviours

Real applications rarely search everything. You filter: this tenant, this category, this year or later. Spring AI gives you one filter expression language for all four stores, and I used category == 'c3' && year >= 2022, which matches 1,328 of the 30,000 documents (4.4 percent). The truth is recomputed over only the matching documents. Results, from 03-filtered.txt:
filter: category == 'c3' && year >= 2022  (1328 of 30000 documents match, 4.4%)
store           recall@10    p50 ms    p95 ms  mean hits
pgvector            0.167      10.0      15.7        1.8
redis               1.000       1.6       2.4       10.0
qdrant              1.000      27.6      38.9       10.0
elasticsearch       0.544       3.1       7.7       10.0
pgvector0.17Redis1.00Qdrant1.00Elasticsearch0.54filtered recall@10 on library defaults (higher is better)
Three different things are going on, and only one of them is visible in the recall column. Redis and Qdrant return all ten correct documents, but Qdrant takes 27.6 ms to do it, about 17 times Redis. Elasticsearch’s recall is its quantisation again (it filters correctly; 0.544 is the same defect as before). And pgvector returns two documents, not ten — its mean hit count is 1.8. A query that asked for ten gets two, with no error.

pgvector: the index finds 40 candidates, then the filter throws most of them away

HNSW index scan returns the 40 nearest of 30,000 (hnsw.ef_search = 40)WHERE category = ‘c3’ AND year >= 2022 — applied after the scan2 rows returned40 candidates × 4.4% matching ≈ 1.8 expected. Measured mean: 1.8 of 10 requested.
The picture is the whole mechanism. pgvector’s HNSW index knows nothing about your WHERE clause. It returns the 40 nearest vectors in the whole table (hnsw.ef_search defaults to 40), and PostgreSQL then applies the filter to those 40. If 4.4 percent of rows match, about 1.8 of the 40 survive, and that is almost exactly what was measured. Raising ef_search to 200 did nothing useful; the count only recovered at 1,000, at a cost in latency. pgvector 0.8 added the proper fix, iterative scan, which keeps walking the index until it has enough matching rows (07-pgvector-ef-search.txt):
ef_search  iterative_scan  mean hits    p50 ms  filt recall
40         off                   1.8       1.2        0.167
200        off                   2.0       1.7        0.198
1000       off                  10.0      28.2        1.000
40         relaxed_order        10.0       7.7        0.611
40         strict_order         10.0       9.7        0.281
With hnsw.iterative_scan = relaxed_order the query returns all ten for 7.7 ms, but only 61 percent of them are the true nearest ten, because the scan stops as soon as it has ten matches rather than the best ten. strict_order returned ten results too and scored lower still on this data (0.281), which I did not expect and did not chase. For a filter this selective, the honest options are a bigger ef_search (correct, slow), iterative scan (fast, approximate) or a partial index or partition on the filter column so the index only contains matching rows (the last I did not test).
Check your pgvector version. The Ubuntu 24.04 package is 0.6.0, which has no iterative scan, so the setting does not exist there. I built 0.8.7 from its tag. Spring AI sets none of these values for you; they are session settings (SET hnsw.ef_search), so they have to be applied on the connection that runs the query, for example through a connection-init SQL on your pool.

Qdrant: right answers, slowly, until you index the field

Qdrant’s filtered queries were exact but took 27.6 ms because Spring AI creates the collection without payload indexes. My reading is that the filter then has to be checked against each candidate’s stored payload during the graph walk; what I measured is the before and after. Creating an index on each filtered field is a REST call per field. The result (08-qdrant-payload-index-and-threshold.txt):
filtered query, no payload index:    recall 1.000  p50 26.2 ms  p95 37.2 ms
filtered query, keyword+integer index: recall 1.000  p50 1.5 ms  p95 3.2 ms
Same recall, about 17 times faster. That belongs in your deployment script for every field you filter on.

Redis: filter on a field you did not declare, and it refuses

Redis builds a schema, so the metadata fields you want to filter on have to be declared in the builder (MetadataField.tag("category"), MetadataField.numeric("year")). I built a store without declaring them and filtered anyway (09-redis-undeclared-metadata.txt):
metadata fields NOT declared; filter category == 'c3' threw IllegalArgumentException
  Not allowed filter identifier name: category
This is the friendliest failure in the article: a clear exception at query time instead of a wrong result. Declare the fields you filter on up front: the schema is built when the index is created.
Going deeper: the Qdrant indexing threshold
Qdrant only builds an HNSW graph for a segment once its vectors pass indexing_threshold, which the live collection reported as 10,000 KB. Below that it searches the segment exactly, which is fast and perfect for small data but not what you are benchmarking. I loaded 10,000 documents into a second collection and compared it with the 30,000-document one (from 08-qdrant-payload-index-and-threshold.txt):
collection vs_small  points_count=10000  segments_count=2   indexed_vectors_count=0      indexing_threshold=10000 KB
collection vs_bench  points_count=30000  segments_count=2   indexed_vectors_count=30000  indexing_threshold=10000 KB
The 10,000-document collection holds 15,000 KB of vectors, which is above 10,000 KB, yet it has indexed_vectors_count=0. The collection has two segments and the threshold applies to each segment, so each holds about 7,500 KB and neither is indexed; the 30,000-document collection has about 23,000 KB per segment and both are. I derived the per-segment reading from these numbers rather than from reading Qdrant’s source, so treat it as the explanation that fits the evidence. The practical point stands: a small Qdrant collection is silently searched by brute force, which is a good property and a misleading benchmark.

What each one costs to keep running

Speed and recall are what you measure first; memory and disk are what you pay for. After loading and querying, I read each service’s proportional memory use from /proc and asked each store how large the loaded data is (11-ops-footprint.txt):
30000 x 384 float32 = 46 MB of raw vectors. PSS = proportional set size of the service's processes, read from /proc, after loading and querying.
store             PSS MB   size of the loaded data, as the store reports it
pgvector             153   table+toast 59 MB, hnsw index 59 MB
redis                422   used_memory 434 MB for 30000 keys (everything lives in RAM); one document = 10600 bytes; vector_index_sz_mb=55.4
qdrant               172   storage dir 144 MB
elasticsearch       1550   index store.size/docs.count [{"store.size":"54.3mb","docs.count":"30000"}] (JVM heap fixed at -Xmx1g)
Elasticsearch54 MBpgvector118 MBQdrant144 MBRedis (RAM)434 MBsize of the loaded data in MB; raw float32 vectors are 46 MB. Redis is memory, the others are disk
The raw vectors are 46 MB, and no store stores just that. Elasticsearch’s 54 MB is the smallest — and it is not the quantised mapping that makes it small: the unquantised hnsw index I built for the previous experiment is the same size (the store column in the Elasticsearch table above), because quantisation keeps the full vectors for re-scoring and adds compressed copies. What it saves is memory at search time, which this measurement does not capture. pgvector keeps the rows and the index at about 59 MB each. Qdrant uses 144 MB of disk. Redis is the outlier: 434 MB of RAM for 30,000 documents, about 10.6 KB per document for a vector that is 1.5 KB raw. I did not trace why (Spring AI’s Redis store keeps each document as JSON and I suspect the vector is stored as text digits, but I did not verify that), so treat the cause as unconfirmed and the number as measured. Because Redis holds everything in memory, that is the number that sets your instance size. Process memory tells a second story: with a 1 GB heap, Elasticsearch’s process was about 1.5 GB, against 150 to 175 MB for pgvector and Qdrant.

Elasticsearch will refuse to store your data on a nearly full disk, and say so only in its log

One failure from the build is worth keeping, because it looks like a bug in your code. Midway through, the benchmark’s bulk insert failed with primary shard is not active Timeout: [1m], and the cluster went red. Nothing in the application was wrong. The sandbox disk had crossed 90 percent used, and Elasticsearch’s default disk watermark refuses to allocate shards on a node that full. The captured evidence is 13-es-disk-watermark.txt:
$ curl -s localhost:9200/_cluster/allocation/explain   (index, shard and decider lines)
index: vs_es_int8_hnsw | can_allocate: no | unassigned reason: INDEX_CREATED
disk_threshold -> NO | the node is above the high watermark cluster setting [cluster.routing.allocation.disk.watermark.high=90%], having less than the minimum required [25.1gb] free space, actual free: [25.1gb], actual used: [90%]
$ grep DiskThresholdMonitor eslogs/elasticsearch.log | tail -1
[WARN ][o.e.c.r.a.DiskThresholdMonitor] [vm] high disk watermark [90%] exceeded on [UYHyLZrWSZSD9gDv_TWuGw][vm][/tmp/vs-run/esdata] free: 25.1gb[9.9%], shards will be relocated away from this node; currently relocating away shards totalling [0] bytes; the node is expected to continue to exceed the high disk watermark when these 
The fingerprint. A fresh index whose shards stay UNASSIGNED with reason INDEX_CREATED, a bulk that times out after a minute, and a single WARN line about the high disk watermark in the Elasticsearch log. Free disk space, or in a throwaway environment set cluster.routing.allocation.disk.threshold_enabled to false, which is what the module’s service script now does. Do not do that in production.
Going deeper: running the four stores without Docker
The build environment had no Docker, so the module’s services-up.sh starts each store from its own distribution: PostgreSQL 16 and the pgvector build from the distribution package plus a source build of tag v0.8.7; Redis Stack from the vendor’s tarball (the Ubuntu 22.04 build, because the 20.04 one links against an OpenSSL the machine lacks); Qdrant’s single static binary; and Elasticsearch’s tarball, which refuses to run as root, so the script creates an es user. The four need roughly 2.5 GB of RAM together. If you have Docker, the official images are simpler and are what you would use in CI.

So which one should you pick?

The measurements do not crown a winner; they say which default to distrust in each. The table puts the choice in terms of what you already run and what you can tolerate.
If you…PickThen do this on day one
already run PostgreSQL and have millions, not billions, of vectorspgvectorUse 0.8 or later; raise ef_search and turn on iterative scan for filtered queries; check mean hit count, not just latency
already run Redis and the vectors fit in RAMRedisSet hnswEfRuntime (the default of 10 gave 0.62 recall here); declare every filter field; budget ten times the raw vector size
want a dedicated vector database with good filteringQdrantCreate payload indexes on every filtered field; know that small collections are searched exactly
already run Elasticsearch for text search and want one engine for bothElasticsearchCreate the index yourself and choose index_options.type; measure recall on your own embeddings; watch the disk watermark
Notice that no row says “pick the fastest”. On defaults the fastest store (Redis) was among the least accurate and the most accurate (Qdrant) was among the slowest to filter. Once each is tuned the four are within a few milliseconds of each other on 30,000 documents, which is a small dataset by vector-database standards. The decision is usually made by what you already operate, because a fifth system to monitor costs more than a few milliseconds.
Should you add a vector database at all? If your documents are already in PostgreSQL and the corpus is under a few million vectors, pgvector in the database you have is almost always the right first answer: one backup, one access-control model, transactional deletes, and joins between vectors and your other tables. Add a dedicated store when you can show, with a measurement on your own data, that pgvector cannot meet your latency or recall at your size — not before. What this article did not test: concurrent query load, datasets beyond 30,000 documents, replication and failover, hybrid text-plus-vector search, real embeddings from any model, or any managed cloud offering. All timings are from one shared 2-vCPU machine with the four services running side by side, so absolute milliseconds will differ on yours; the recall figures and the shape of each comparison are what to take away.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.