Files

vector-stores

Companion code for Choosing a Vector Store for Spring AI: pgvector vs Redis vs Qdrant vs Elasticsearch, part of the Spring AI series on ankurm.com.

The same 30,000-document, 384-dimension dataset goes through Spring AI's VectorStore API into four real stores, each left on the library's defaults. Ingest time, recall@10 against an exact brute-force answer, latency, metadata filtering and the cost of running each one are measured by one test class, and every figure in the article is quoted from a file in output/.

The dataset is synthetic (clustered unit vectors, looked up by LookupEmbeddingModel), so no API key is needed and all four stores receive identical vectors. It measures the stores, not an embedding model: recall on real embeddings will differ, particularly for Elasticsearch's quantised default.

Versions

Component Version
Spring Boot 4.1.1
Spring AI 2.0.1
Java 25 (Temurin 25.0.4.1)
PostgreSQL / pgvector 16.15 / 0.8.7, built from the v0.8.7 tag (the Ubuntu package is 0.6.0, which has no iterative scan)
Redis Stack 7.4.0-v8 tarball (Redis 7.4.7, RediSearch 2.10.20)
Qdrant 1.19.2 (single binary; Java client 1.18.0)
Elasticsearch 9.5.5 single node, security off (Java client 9.4.5 from the Spring AI BOM)
Jedis 7.4.1

Quickstart (no Docker)

scripts/services-up.sh     # pgvector (apt + source build), Redis Stack, Qdrant, Elasticsearch -- see the article for the download URLs
scripts/run-all.sh         # runs the 11 tests and regenerates output/01 .. 12

A store whose port is closed is skipped, not failed. The four services together need about 2.5 GB of RAM.

What's here

File What it shows
Dataset.java The synthetic vectors, metadata and the exact top-k used as the answer key
LookupEmbeddingModel.java An EmbeddingModel that looks vectors up, so every store gets the same ones
StoreFactory.java The four stores built by hand with library defaults, and why afterPropertiesSet() is called
StoreComparisonTest.java Every measurement; each test writes its own transcript
src/broken/Redis1xStyle.java Not compiled by the build; scripts/capture-1x-compile.sh compiles it against 2.0.1

Output files

File Written by
01-ingest.txt .. 03-filtered.txt ingestTimes, unfilteredRecallAndLatency, filteredRecallAndLatency
04-defaults-that-matter.txt settingsThatSurprise (pgvector ef_search, the Elasticsearch mapping, Qdrant HNSW config)
05-elasticsearch-num-candidates.txt elasticsearchCandidatesNotQuantisation
06-redis-ef-runtime.txt redisEfRuntime
07-pgvector-ef-search.txt pgvectorEfSearchAndFilters
08-qdrant-payload-index-and-threshold.txt qdrantPayloadIndexAndIndexingThreshold
09-redis-undeclared-metadata.txt redisFilterWithoutDeclaredSchema
10-elasticsearch-index-options.txt elasticsearchIndexOptions
11-ops-footprint.txt operationalFootprint
12-redis-1x-compile.txt scripts/capture-1x-compile.sh
13-es-disk-watermark.txt captured once by hand: Elasticsearch refusing to allocate shards on a disk above 90% (not regenerated by run-all.sh)

Timings drift between runs (this is a 2-vCPU sandbox with all four services sharing it); the recall figures and the shape of every comparison do not.

Not covered

Concurrent query load, very large datasets, replication, hybrid (text + vector) search, and any embedding model's real recall.