VectorStore, in front of a long list of them. The interface makes switching look free: change a dependency, change a few properties, and the same add() and similaritySearch() calls run against a different database.
It is not free, and the part that is not free is quiet. Every store has defaults, the defaults differ, and the differences show up as answers that are slightly wrong rather than as errors. This article puts the same 30,000 documents into four stores — pgvector, Redis, Qdrant and Elasticsearch — through Spring AI 2.0’s VectorStore API, leaves every store on the library’s defaults, and measures how fast each one ingests, how many of the true nearest neighbours each one actually returns, how each behaves when you add a metadata filter, and what each costs to keep running. Then it shows, for each surprise, the single setting that fixes it. Depth is in expandable sections; read straight through or open only what you need.
Versions this was written and run against. Spring Boot 4.1.1, Spring AI 2.0.1 and Java 25. The stores were real processes on one 2-vCPU machine, no Docker: PostgreSQL 16.15 with pgvector 0.8.7 (built from itsv0.8.7tag), Redis Stack 7.4.0-v8 (Redis 7.4.7 with RediSearch 2.10.20), Qdrant 1.19.2 and Elasticsearch 9.5.5 (single node, security off). All the code is in thevector-storesmodule of asmhatre/spring-ai; every console block below is quoted from a file under itsoutput/directory, written by a test that asserts the same numbers. The data is synthetic (clustered random vectors, no embedding model), so the numbers compare the stores; they say nothing about how any embedding model performs, and recall on real embeddings will differ — most of all for Elasticsearch’s quantised default.
A vector store answers one question: which of these is closest to that?
A language model can turn a piece of text into a list of numbers, usually a few hundred of them, called an embedding. The useful property is that texts with similar meaning get lists that point in similar directions. “How do I reset my password” and “I forgot my login” end up close together; “quarterly revenue” ends up far away. A vector store keeps those lists, and when you hand it the embedding of a question it returns the stored items whose lists are nearest. That is retrieval-augmented generation in one sentence, and it is why a store sits behind almost every Spring AI application that talks about your own data. Finding the exact nearest items means comparing the question against every stored vector. That is fine for a thousand documents and too slow for ten million, so real stores build an index that finds almost the nearest ones much faster. The four stores here all use the same family of index, HNSW (a layered graph that you walk towards the answer). “Almost” is the important word, and the rest of the article is about how much “almost” each store gives you by default.EmbeddingModel that looks them up by name (d12 is document 12), so no store can win by being fed better data. The dashed lines are the part Spring AI hides from you, and the part this article opens.
Going deeper: the dataset, and what “recall@10” means here
Recall@10 is the share of the ten truly nearest documents that a store actually returned. The truth is computed by brute force in
Dataset.java: for each of 100 queries, a plain loop scores all 30,000 vectors and keeps the ten best. A store that returns eight of those ten has recall 0.8. The stores are all asked for topK=10 with cosine distance.
The dataset is 30,000 unit-length vectors of 384 dimensions grouped into 40 clusters, with three metadata fields per document: category (eight values, tied to the cluster), year (2015–2025) and region. It is synthetic because there is no embedding API key in the build and a downloaded model would make the repository depend on a model host. The cost of that choice is real: quantisation-based indexes in particular can behave differently on real embeddings, and the article says so where it matters.
LookupEmbeddingModel.java— the stand-in model: it implementsEmbeddingModeland returns the stored vector ford123orq7.StoreComparisonTest.java— every measurement in this article; each test writes its own transcript.
One interface, four stores — and one thing the interface does not do for you
With Spring Boot’s starters you normally never build a store yourself; the auto-configuration does. To put four stores in one program, the module builds each by hand with its builder, using the library’s defaults wherever the builder lets it (the file isStoreFactory.java). Here is the pgvector one — the other three are the same shape.
public static VectorStore pgvector(EmbeddingModel em) {
JdbcTemplate jdbc = pgJdbc();
jdbc.execute("CREATE EXTENSION IF NOT EXISTS vector");
return init(PgVectorStore.builder(jdbc, em)
.vectorTableName(NAME)
.dimensions(Dataset.DIMS)
.distanceType(PgVectorStore.PgDistanceType.COSINE_DISTANCE)
.indexType(PgVectorStore.PgIndexType.HNSW)
.removeExistingVectorStoreTable(true)
.initializeSchema(true)
.build());
}
Building by hand exposed the first trap, which would also hit you in a unit test that constructs a store directly. The first run failed on the very first insert:
The table does not exist yet.build()returns a store but creates nothing. The pgvector table, the Redis index, the Qdrant collection and the Elasticsearch index are all created inafterPropertiesSet(), which Spring calls for a bean and nobody calls for a builder. My first run died withERROR: relation "public.vs_bench" does not exist. The fix is theinit()helper inStoreFactory.java, shown next, which calls it explicitly. In a normal Boot application you will never see this.
static VectorStore init(VectorStore store) {
try {
((org.springframework.beans.factory.InitializingBean) store).afterPropertiesSet();
} catch (Exception e) {
throw new IllegalStateException(e);
}
return store;
}
If you are coming from Spring AI 1.x, two more things changed on the way in, both in the Redis store. I compiled the 1.x way of wiring it against 2.0.1 and kept the compiler’s own words:
$ javac -cp <spring-ai 2.0.1 classpath> src/broken/Redis1xStyle.java
src/broken/Redis1xStyle.java:11: error: cannot find symbol
.vectorAlgorithm(RedisVectorStore.Algorithm.HSNW)
^
symbol: variable HSNW
location: class Algorithm
src/broken/Redis1xStyle.java:10: error: incompatible types: JedisPooled cannot be converted to RedisClient
return RedisVectorStore.builder(jedis, em)
The builder now takes Jedis 7’s RedisClient rather than JedisPooled, and the enum constant that was spelled HSNW in 1.x is now HNSW. The file that produces this message is Redis1xStyle.java; it is deliberately outside the build.
Going deeper: why not just use the four starters?
Each store has a starter (
spring-ai-starter-vector-store-pgvector and so on). Put all four on one classpath and Boot would try to create four VectorStore beans in one context. The module depends on the store artifacts (spring-ai-pgvector-store, spring-ai-redis-store, spring-ai-qdrant-store, spring-ai-elasticsearch-store) directly and builds each one itself, so the comparison is in one process. In your application you want exactly one starter. The versions come from the Spring AI BOM: Jedis 7.4.1, the Qdrant Java client 1.18.0 and the Elasticsearch Java client 9.4.5 (note the client is one minor version behind the 9.5.5 server I ran; it worked).
pom.xml— the dependencies, with the reason for not using starters in a comment.- Spring AI reference: Vector Databases
- Spring AI 1.x to 2.0 migration guide on this site.
Getting the data in: the fastest and the slowest differ by more than an order of magnitude
The first measurement is simply how long it takes to add 30,000 documents in batches of 1,000. Embeddings come from the lookup, so this is time spent in the store and its index, not in an embedding API. Here is the transcript from01-ingest.txt:
dataset: 30000 documents x 384 dims, 40 clusters, batches of 1000; embeddings are looked up, so this is store time only
store seconds docs/second
pgvector 46.5 645
redis 9.5 3158
qdrant 1.7 17382
elasticsearch 9.9 3039
Qdrant took 1.7 seconds and pgvector 46.5, a ratio of about 27 to one. Redis and Elasticsearch landed together at about 9.5 seconds. The likely reason pgvector is slowest is that its HNSW index is maintained inside each insert, row by row, from the first row on — I did not isolate that. Spring AI’s store creates the index up front, so you cannot try the usual alternative (load first, build the index after) through the store; you would do that in plain SQL.
“Ingest finished” is not “searchable”. Elasticsearch needs a refresh before new documents appear, Qdrant runs an optimizer pass in the background, and Redis drains an indexing queue. A benchmark, or an integration test, that searches the instantadd()returns can see a partial index. The test waits for each one (a_refresh, a green collection status,percent_indexedof 1) before it measures anything.
Going deeper: ingest in production
These are single-writer, single-node numbers on a shared 2-vCPU box, so read them as ratios, not as capacity. Two things change the picture in production. Embedding time usually dwarfs store time (a hosted embedding API at a few hundred documents per second is slower than any of these stores), so ingest speed rarely decides the choice. If pgvector’s per-row index maintenance is what makes it slow (unverified here), loading in bulk and indexing afterwards in plain SQL is the usual remedy for a large migration.
Finding the nearest ten: two of four stores miss half of them by default
Now the question that matters. 100 queries, ten results each, against the exact answer. Nothing is tuned: this is what you get by following each store’s reference page. The transcript is02-unfiltered.txt:
100 queries after 20 warm-up, topK=10, cosine, each store's library defaults
store recall@10 p50 ms p95 ms mean hits
pgvector 0.951 9.5 16.4 10.0
redis 0.616 0.7 1.1 10.0
qdrant 0.999 2.8 6.8 10.0
elasticsearch 0.491 5.4 20.5 10.0
Redis: one runtime setting, and recall goes from 0.62 to 0.96
Redis’s HNSW index has a query-time knob,EF_RUNTIME, which is how many candidates the graph walk keeps before it stops. Left unset, Redis uses 10 — the same as the topK of 10 I asked for, which leaves no slack for the walk to recover from a wrong turn. Spring AI exposes it as hnswEfRuntime(...) on the builder. I rebuilt the index with three settings and loaded the same documents each time (06-redis-ef-runtime.txt):
Redis HNSW EF_RUNTIME: the library default (not set by Spring AI) against explicit values; index rebuilt and reloaded each time
ef_runtime recall@10 p50 ms filt recall filt p50 ms
default 0.616 0.6 1.000 1.6
50 0.962 0.6 1.000 1.3
200 1.000 0.8 1.000 1.3
At 50, recall is 0.962 for no measurable cost in latency; at 200 it is 1.000. The fix is one line, .hnswEfRuntime(100), and I would set it explicitly in any Redis-backed application rather than inherit 10.
Elasticsearch: the default mapping is quantised, and on this data it costs half the recall
Elasticsearch’s result was the surprise, and the first explanation I wrote down was wrong, so here is the path. The Spring AI store asks fork=10 with num_candidates set to 1.5 times topK, which is 15 — I read that straight from the store’s bytecode (ldc2_w 1.5d, dmul, d2i). Fifteen candidates for ten results looks stingy, so I assumed that was the cause and tested it by sending the same query directly to Elasticsearch with num_candidates of 15, 50, 100 and 500 (05-elasticsearch-num-candidates.txt):
same index, same 100 queries, raw _search with the num_candidates varied:
num_candidates recall@10 p50 ms
15 0.491 4.5
50 0.489 3.9
100 0.490 3.8
500 0.490 4.0
Recall did not move. Whatever limits it, it is not the candidate count. The next suspect was the index itself. The mapping the store creates, captured from the live index in 04-defaults-that-matter.txt, is:
elasticsearch embedding mapping: "embedding":{"type":"dense_vector","dims":384,"index":true,"similarity":"cosine","index_options":{"type":"bbq_hnsw","m":16,"ef_construction":100,"rescore_vector":{"oversample":3.0}}}
bbq_hnsw is what Elasticsearch applied here: its documentation makes it the default for vectors of 384 dimensions or more. It compresses each vector to about one bit per dimension (“better binary quantization”), searches the small copies, then re-scores the top candidates with the full vectors (oversample: 3.0). It saves a great deal of memory. To see what it costs here, I created two more indexes by hand that differ from the default in exactly one field, index_options.type, loaded the same documents through the same Spring AI store, and measured (10-elasticsearch-index-options.txt):
type recall@10 p50 ms filt recall filt p50 ms store
bbq_hnsw* 0.491 3.6 0.544 3.0 54.4mb
int8_hnsw 0.952 3.0 0.974 3.1 64.1mb
hnsw 0.995 3.0 1.000 2.7 52.6mb
* default; its full mapping is in 04-defaults-that-matter.txt
Same data, same store class, same queries: bbq_hnsw gives 0.491, 8-bit quantisation (int8_hnsw) gives 0.952, and unquantised hnsw gives 0.995. The index type is the cause. Two honesty points belong here. Elasticsearch builds its graph non-deterministically, so the default row moved between 0.47 and 0.50 over the runs I did; the gap to the others is far larger than that wobble. And this data is synthetic, a stress test for one-bit quantisation, because my vectors are random points around cluster centres with no fine structure to preserve. Elastic’s claim that BBQ holds recall on real embeddings is not something this experiment can confirm or refute. What it does show is that the default is a decision made for you.
What to do. Spring AI’sElasticsearchVectorStoreOptionshas no field forindex_options, so you cannot choose the type through the store. Create the index yourself with the mapping you want (the test shows thePUT) and build the store withinitializeSchema(false). Then measure recall on a sample of your embeddings before you trust either choice.
Going deeper: HNSW in two minutes, and the three numbers that matter
HNSW stores each vector as a node in a graph with several layers. A search enters at the sparse top layer, hops to the neighbour closest to the query, drops down a layer, and repeats until the dense bottom layer, where it keeps exploring until it has looked at
ef candidates. Three numbers control the trade between speed and accuracy: m (links per node, fixed at build time), ef_construction (candidates considered while building) and ef at query time (called ef_search in pgvector, EF_RUNTIME in Redis, ef/hnsw_ef in Qdrant, and num_candidates in Elasticsearch). A larger query-time ef costs more time and finds more of the true neighbours.
The defaults in this run were: pgvector m=16, ef_construction=64, ef_search=40; Redis EF_RUNTIME=10 (my reading of the library behaviour, confirmed by the recall change above); Qdrant m=16, ef_construct=100 (read from the live collection); Elasticsearch m=16, ef_construction=100.
04-defaults-that-matter.txt— the pgvector index definition, the Elasticsearch mapping and the Qdrant HNSW configuration, captured from the live services.- Elasticsearch: dense_vector and index_options
- Redis: vector fields and HNSW parameters
Adding a metadata filter: the same query, four different behaviours
Real applications rarely search everything. You filter: this tenant, this category, this year or later. Spring AI gives you one filter expression language for all four stores, and I usedcategory == 'c3' && year >= 2022, which matches 1,328 of the 30,000 documents (4.4 percent). The truth is recomputed over only the matching documents. Results, from 03-filtered.txt:
filter: category == 'c3' && year >= 2022 (1328 of 30000 documents match, 4.4%)
store recall@10 p50 ms p95 ms mean hits
pgvector 0.167 10.0 15.7 1.8
redis 1.000 1.6 2.4 10.0
qdrant 1.000 27.6 38.9 10.0
elasticsearch 0.544 3.1 7.7 10.0
pgvector: the index finds 40 candidates, then the filter throws most of them away
WHERE clause. It returns the 40 nearest vectors in the whole table (hnsw.ef_search defaults to 40), and PostgreSQL then applies the filter to those 40. If 4.4 percent of rows match, about 1.8 of the 40 survive, and that is almost exactly what was measured. Raising ef_search to 200 did nothing useful; the count only recovered at 1,000, at a cost in latency. pgvector 0.8 added the proper fix, iterative scan, which keeps walking the index until it has enough matching rows (07-pgvector-ef-search.txt):
ef_search iterative_scan mean hits p50 ms filt recall
40 off 1.8 1.2 0.167
200 off 2.0 1.7 0.198
1000 off 10.0 28.2 1.000
40 relaxed_order 10.0 7.7 0.611
40 strict_order 10.0 9.7 0.281
With hnsw.iterative_scan = relaxed_order the query returns all ten for 7.7 ms, but only 61 percent of them are the true nearest ten, because the scan stops as soon as it has ten matches rather than the best ten. strict_order returned ten results too and scored lower still on this data (0.281), which I did not expect and did not chase. For a filter this selective, the honest options are a bigger ef_search (correct, slow), iterative scan (fast, approximate) or a partial index or partition on the filter column so the index only contains matching rows (the last I did not test).
Check your pgvector version. The Ubuntu 24.04 package is 0.6.0, which has no iterative scan, so the setting does not exist there. I built 0.8.7 from its tag. Spring AI sets none of these values for you; they are session settings (SET hnsw.ef_search), so they have to be applied on the connection that runs the query, for example through a connection-init SQL on your pool.
Qdrant: right answers, slowly, until you index the field
Qdrant’s filtered queries were exact but took 27.6 ms because Spring AI creates the collection without payload indexes. My reading is that the filter then has to be checked against each candidate’s stored payload during the graph walk; what I measured is the before and after. Creating an index on each filtered field is a REST call per field. The result (08-qdrant-payload-index-and-threshold.txt):
filtered query, no payload index: recall 1.000 p50 26.2 ms p95 37.2 ms
filtered query, keyword+integer index: recall 1.000 p50 1.5 ms p95 3.2 ms
Same recall, about 17 times faster. That belongs in your deployment script for every field you filter on.
Redis: filter on a field you did not declare, and it refuses
Redis builds a schema, so the metadata fields you want to filter on have to be declared in the builder (MetadataField.tag("category"), MetadataField.numeric("year")). I built a store without declaring them and filtered anyway (09-redis-undeclared-metadata.txt):
metadata fields NOT declared; filter category == 'c3' threw IllegalArgumentException
Not allowed filter identifier name: category
This is the friendliest failure in the article: a clear exception at query time instead of a wrong result. Declare the fields you filter on up front: the schema is built when the index is created.
Going deeper: the Qdrant indexing threshold
Qdrant only builds an HNSW graph for a segment once its vectors pass
indexing_threshold, which the live collection reported as 10,000 KB. Below that it searches the segment exactly, which is fast and perfect for small data but not what you are benchmarking. I loaded 10,000 documents into a second collection and compared it with the 30,000-document one (from 08-qdrant-payload-index-and-threshold.txt):
collection vs_small points_count=10000 segments_count=2 indexed_vectors_count=0 indexing_threshold=10000 KB
collection vs_bench points_count=30000 segments_count=2 indexed_vectors_count=30000 indexing_threshold=10000 KB
The 10,000-document collection holds 15,000 KB of vectors, which is above 10,000 KB, yet it has indexed_vectors_count=0. The collection has two segments and the threshold applies to each segment, so each holds about 7,500 KB and neither is indexed; the 30,000-document collection has about 23,000 KB per segment and both are. I derived the per-segment reading from these numbers rather than from reading Qdrant’s source, so treat it as the explanation that fits the evidence. The practical point stands: a small Qdrant collection is silently searched by brute force, which is a good property and a misleading benchmark.
What each one costs to keep running
Speed and recall are what you measure first; memory and disk are what you pay for. After loading and querying, I read each service’s proportional memory use from/proc and asked each store how large the loaded data is (11-ops-footprint.txt):
30000 x 384 float32 = 46 MB of raw vectors. PSS = proportional set size of the service's processes, read from /proc, after loading and querying.
store PSS MB size of the loaded data, as the store reports it
pgvector 153 table+toast 59 MB, hnsw index 59 MB
redis 422 used_memory 434 MB for 30000 keys (everything lives in RAM); one document = 10600 bytes; vector_index_sz_mb=55.4
qdrant 172 storage dir 144 MB
elasticsearch 1550 index store.size/docs.count [{"store.size":"54.3mb","docs.count":"30000"}] (JVM heap fixed at -Xmx1g)
hnsw index I built for the previous experiment is the same size (the store column in the Elasticsearch table above), because quantisation keeps the full vectors for re-scoring and adds compressed copies. What it saves is memory at search time, which this measurement does not capture. pgvector keeps the rows and the index at about 59 MB each. Qdrant uses 144 MB of disk. Redis is the outlier: 434 MB of RAM for 30,000 documents, about 10.6 KB per document for a vector that is 1.5 KB raw. I did not trace why (Spring AI’s Redis store keeps each document as JSON and I suspect the vector is stored as text digits, but I did not verify that), so treat the cause as unconfirmed and the number as measured. Because Redis holds everything in memory, that is the number that sets your instance size. Process memory tells a second story: with a 1 GB heap, Elasticsearch’s process was about 1.5 GB, against 150 to 175 MB for pgvector and Qdrant.
Elasticsearch will refuse to store your data on a nearly full disk, and say so only in its log
One failure from the build is worth keeping, because it looks like a bug in your code. Midway through, the benchmark’s bulk insert failed withprimary shard is not active Timeout: [1m], and the cluster went red. Nothing in the application was wrong. The sandbox disk had crossed 90 percent used, and Elasticsearch’s default disk watermark refuses to allocate shards on a node that full. The captured evidence is 13-es-disk-watermark.txt:
$ curl -s localhost:9200/_cluster/allocation/explain (index, shard and decider lines)
index: vs_es_int8_hnsw | can_allocate: no | unassigned reason: INDEX_CREATED
disk_threshold -> NO | the node is above the high watermark cluster setting [cluster.routing.allocation.disk.watermark.high=90%], having less than the minimum required [25.1gb] free space, actual free: [25.1gb], actual used: [90%]
$ grep DiskThresholdMonitor eslogs/elasticsearch.log | tail -1
[WARN ][o.e.c.r.a.DiskThresholdMonitor] [vm] high disk watermark [90%] exceeded on [UYHyLZrWSZSD9gDv_TWuGw][vm][/tmp/vs-run/esdata] free: 25.1gb[9.9%], shards will be relocated away from this node; currently relocating away shards totalling [0] bytes; the node is expected to continue to exceed the high disk watermark when these
The fingerprint. A fresh index whose shards stayUNASSIGNEDwith reasonINDEX_CREATED, a bulk that times out after a minute, and a single WARN line about the high disk watermark in the Elasticsearch log. Free disk space, or in a throwaway environment setcluster.routing.allocation.disk.threshold_enabledto false, which is what the module’s service script now does. Do not do that in production.
Going deeper: running the four stores without Docker
The build environment had no Docker, so the module’s
services-up.sh starts each store from its own distribution: PostgreSQL 16 and the pgvector build from the distribution package plus a source build of tag v0.8.7; Redis Stack from the vendor’s tarball (the Ubuntu 22.04 build, because the 20.04 one links against an OpenSSL the machine lacks); Qdrant’s single static binary; and Elasticsearch’s tarball, which refuses to run as root, so the script creates an es user. The four need roughly 2.5 GB of RAM together. If you have Docker, the official images are simpler and are what you would use in CI.
services-up.shrun-all.sh— regenerates every file inoutput/.
So which one should you pick?
The measurements do not crown a winner; they say which default to distrust in each. The table puts the choice in terms of what you already run and what you can tolerate.| If you… | Pick | Then do this on day one |
|---|---|---|
| already run PostgreSQL and have millions, not billions, of vectors | pgvector | Use 0.8 or later; raise ef_search and turn on iterative scan for filtered queries; check mean hit count, not just latency |
| already run Redis and the vectors fit in RAM | Redis | Set hnswEfRuntime (the default of 10 gave 0.62 recall here); declare every filter field; budget ten times the raw vector size |
| want a dedicated vector database with good filtering | Qdrant | Create payload indexes on every filtered field; know that small collections are searched exactly |
| already run Elasticsearch for text search and want one engine for both | Elasticsearch | Create the index yourself and choose index_options.type; measure recall on your own embeddings; watch the disk watermark |
Should you add a vector database at all? If your documents are already in PostgreSQL and the corpus is under a few million vectors, pgvector in the database you have is almost always the right first answer: one backup, one access-control model, transactional deletes, and joins between vectors and your other tables. Add a dedicated store when you can show, with a measurement on your own data, that pgvector cannot meet your latency or recall at your size — not before. What this article did not test: concurrent query load, datasets beyond 30,000 documents, replication and failover, hybrid text-plus-vector search, real embeddings from any model, or any managed cloud offering. All timings are from one shared 2-vCPU machine with the four services running side by side, so absolute milliseconds will differ on yours; the recall figures and the shape of each comparison are what to take away.
Further reading
- The code for this article: the
vector-storesmodule of asmhatre/spring-ai, with all the transcripts underoutput/and its README. - Production-grade RAG with Spring AI and the complete RAG example — what sits on top of the store you pick here.
- Chat memory in Spring AI 2.0: JDBC, Redis and windowed conversations — the same PostgreSQL and Redis, used for conversation history.
- Spring AI 2.0 ChatClient on Spring Boot 4.1 and the 1.x to 2.0 migration guide.
- Spring AI reference: Vector Databases
- pgvector, Qdrant documentation, Redis vector search and Elasticsearch dense_vector.
- Malkov and Yashunin, Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs, the HNSW paper.
No Comments yet!