diff --git a/README.md b/README.md index f9588c3..ded338e 100644 --- a/README.md +++ b/README.md @@ -14,5 +14,6 @@ Runnable companion code for the Spring AI articles on [ankurm.com](https://ankur | [`ollama-local/`](ollama-local) | Chat and embeddings against a real local `qwen2.5:0.5b`/`all-minilm`, no API key, driven by a Testcontainers-managed Ollama container started from a baked image; a confirmed model unload via `keep_alive: 0` and `/api/ps`, not a scripted model anywhere. Spring Boot 4.1.1, Spring AI 2.0.1, Testcontainers 2.0.5, Java 25. | [Run LLMs Locally with Spring AI and Ollama](https://ankurm.com/spring-ai-2-0-ollama-local/) | | [`chat-memory/`](chat-memory) | `MessageChatMemoryAdvisor`, `MessageWindowChatMemory`, the JDBC and Redis `ChatMemoryRepository`, per-user conversation IDs and a token-budget memory of our own, with the traps reproduced against a real PostgreSQL 16 and Redis Stack: a 36-character `conversation_id`, tool messages dropped on save, concurrent writers, a 1.x table under the 2.0 repository, and a Redis repository that silently steps aside for a custom `ChatMemory`. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. | [Chat Memory in Spring AI 2.0: JDBC, Redis and Windowed Conversations](https://ankurm.com/spring-ai-2-0-chat-memory-jdbc-redis-windowed-conversations/) | | [`advisors/`](advisors) | Three custom advisors -- a logger, a PII redactor (with a stream-safe restore) and a per-request / per-user token budget -- and tests for how the chain is ordered, what `BaseAdvisor` does on a stream, where an advisor sits relative to memory and the tool loop, and what a refusal looks like on a call, a stream and over HTTP (429). A recording stub model, no live model. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. | [Writing Custom Advisors in Spring AI 2.0: Logging, PII Redaction and Token Budgets](https://ankurm.com/spring-ai-2-0-custom-advisors-logging-pii-redaction-token-budgets/) | +| [`vector-stores/`](vector-stores) | The same 30,000-document dataset behind `VectorStore` on pgvector, Redis, Qdrant and Elasticsearch: ingest time, recall@10, latency, metadata filtering and running cost, with the defaults that cost recall reproduced (Elasticsearch's quantised mapping, Redis `EF_RUNTIME`, pgvector post-filtering, Qdrant payload indexes). Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. | [Choosing a Vector Store for Spring AI](https://ankurm.com/spring-ai-2-0-vector-store-comparison-pgvector-redis-qdrant-elasticsearch/) | Upgrading from Spring AI 1.x: [migration guide](https://ankurm.com/spring-ai-1-to-2-migration-guide/). diff --git a/vector-stores/.gitignore b/vector-stores/.gitignore new file mode 100644 index 0000000..2f7896d --- /dev/null +++ b/vector-stores/.gitignore @@ -0,0 +1 @@ +target/ diff --git a/vector-stores/README.md b/vector-stores/README.md new file mode 100644 index 0000000..2abe232 --- /dev/null +++ b/vector-stores/README.md @@ -0,0 +1,61 @@ +# vector-stores + +Companion code for [Choosing a Vector Store for Spring AI: pgvector vs Redis vs Qdrant vs Elasticsearch](https://ankurm.com/spring-ai-2-0-vector-store-comparison-pgvector-redis-qdrant-elasticsearch/), part of the [Spring AI series](../README.md) on ankurm.com. + +The same 30,000-document, 384-dimension dataset goes through Spring AI's `VectorStore` API into four real stores, each left on the library's defaults. Ingest time, recall@10 against an exact brute-force answer, latency, metadata filtering and the cost of running each one are measured by one test class, and every figure in the article is quoted from a file in [`output/`](output). + +The dataset is **synthetic** (clustered unit vectors, looked up by [`LookupEmbeddingModel`](src/main/java/com/ankurm/vectorstores/LookupEmbeddingModel.java)), so no API key is needed and all four stores receive identical vectors. It measures the *stores*, not an embedding model: recall on real embeddings will differ, particularly for Elasticsearch's quantised default. + +## Versions + +| Component | Version | +|---|---| +| Spring Boot | 4.1.1 | +| Spring AI | 2.0.1 | +| Java | 25 (Temurin 25.0.4.1) | +| PostgreSQL / pgvector | 16.15 / **0.8.7**, built from the `v0.8.7` tag (the Ubuntu package is 0.6.0, which has no iterative scan) | +| Redis Stack | 7.4.0-v8 tarball (Redis 7.4.7, RediSearch 2.10.20) | +| Qdrant | 1.19.2 (single binary; Java client 1.18.0) | +| Elasticsearch | 9.5.5 single node, security off (Java client 9.4.5 from the Spring AI BOM) | +| Jedis | 7.4.1 | + +## Quickstart (no Docker) + +```bash +scripts/services-up.sh # pgvector (apt + source build), Redis Stack, Qdrant, Elasticsearch -- see the article for the download URLs +scripts/run-all.sh # runs the 11 tests and regenerates output/01 .. 12 +``` + +A store whose port is closed is skipped, not failed. The four services together need about 2.5 GB of RAM. + +## What's here + +| File | What it shows | +|---|---| +| [`Dataset.java`](src/main/java/com/ankurm/vectorstores/Dataset.java) | The synthetic vectors, metadata and the exact top-k used as the answer key | +| [`LookupEmbeddingModel.java`](src/main/java/com/ankurm/vectorstores/LookupEmbeddingModel.java) | An `EmbeddingModel` that looks vectors up, so every store gets the same ones | +| [`StoreFactory.java`](src/main/java/com/ankurm/vectorstores/StoreFactory.java) | The four stores built by hand with library defaults, and why `afterPropertiesSet()` is called | +| [`StoreComparisonTest.java`](src/test/java/com/ankurm/vectorstores/StoreComparisonTest.java) | Every measurement; each test writes its own transcript | +| [`src/broken/Redis1xStyle.java`](src/broken/Redis1xStyle.java) | Not compiled by the build; `scripts/capture-1x-compile.sh` compiles it against 2.0.1 | + +## Output files + +| File | Written by | +|---|---| +| `01-ingest.txt` .. `03-filtered.txt` | `ingestTimes`, `unfilteredRecallAndLatency`, `filteredRecallAndLatency` | +| `04-defaults-that-matter.txt` | `settingsThatSurprise` (pgvector `ef_search`, the Elasticsearch mapping, Qdrant HNSW config) | +| `05-elasticsearch-num-candidates.txt` | `elasticsearchCandidatesNotQuantisation` | +| `06-redis-ef-runtime.txt` | `redisEfRuntime` | +| `07-pgvector-ef-search.txt` | `pgvectorEfSearchAndFilters` | +| `08-qdrant-payload-index-and-threshold.txt` | `qdrantPayloadIndexAndIndexingThreshold` | +| `09-redis-undeclared-metadata.txt` | `redisFilterWithoutDeclaredSchema` | +| `10-elasticsearch-index-options.txt` | `elasticsearchIndexOptions` | +| `11-ops-footprint.txt` | `operationalFootprint` | +| `12-redis-1x-compile.txt` | `scripts/capture-1x-compile.sh` | +| `13-es-disk-watermark.txt` | captured once by hand: Elasticsearch refusing to allocate shards on a disk above 90% (not regenerated by `run-all.sh`) | + +Timings drift between runs (this is a 2-vCPU sandbox with all four services sharing it); the recall figures and the shape of every comparison do not. + +## Not covered + +Concurrent query load, very large datasets, replication, hybrid (text + vector) search, and any embedding model's real recall. diff --git a/vector-stores/output/01-ingest.txt b/vector-stores/output/01-ingest.txt new file mode 100644 index 0000000..59be670 --- /dev/null +++ b/vector-stores/output/01-ingest.txt @@ -0,0 +1,6 @@ +dataset: 30000 documents x 384 dims, 40 clusters, batches of 1000; embeddings are looked up, so this is store time only +store seconds docs/second +pgvector 46.5 645 +redis 9.5 3158 +qdrant 1.7 17382 +elasticsearch 9.9 3039 diff --git a/vector-stores/output/02-unfiltered.txt b/vector-stores/output/02-unfiltered.txt new file mode 100644 index 0000000..2a4015c --- /dev/null +++ b/vector-stores/output/02-unfiltered.txt @@ -0,0 +1,6 @@ +100 queries after 20 warm-up, topK=10, cosine, each store's library defaults +store recall@10 p50 ms p95 ms mean hits +pgvector 0.951 9.5 16.4 10.0 +redis 0.616 0.7 1.1 10.0 +qdrant 0.999 2.8 6.8 10.0 +elasticsearch 0.491 5.4 20.5 10.0 diff --git a/vector-stores/output/03-filtered.txt b/vector-stores/output/03-filtered.txt new file mode 100644 index 0000000..ede490e --- /dev/null +++ b/vector-stores/output/03-filtered.txt @@ -0,0 +1,6 @@ +filter: category == 'c3' && year >= 2022 (1328 of 30000 documents match, 4.4%) +store recall@10 p50 ms p95 ms mean hits +pgvector 0.167 10.0 15.7 1.8 +redis 1.000 1.6 2.4 10.0 +qdrant 1.000 27.6 38.9 10.0 +elasticsearch 0.544 3.1 7.7 10.0 diff --git a/vector-stores/output/04-defaults-that-matter.txt b/vector-stores/output/04-defaults-that-matter.txt new file mode 100644 index 0000000..efc6c0b --- /dev/null +++ b/vector-stores/output/04-defaults-that-matter.txt @@ -0,0 +1,7 @@ +pgvector hnsw.ef_search = 40 +pgvector index: CREATE INDEX vs_bench_index ON public.vs_bench USING hnsw (embedding vector_cosine_ops) +elasticsearch embedding mapping: "embedding":{"type":"dense_vector","dims":384,"index":true,"similarity":"cosine","index_options":{"type":"bbq_hnsw","m":16,"ef_construction":100,"rescore_vector":{"oversample":3.0}}} +qdrant points_count=30000 indexed_vectors_count=30000 +qdrant hnsw_config: "hnsw_config":{"m":16,"ef_construct":100,"full_scan_threshold":10000,"max_indexing_threads":0,"on_disk":false} +qdrant indexing_threshold (KB): 10000 +redis num_docs=30000 percent_indexed=1 diff --git a/vector-stores/output/05-elasticsearch-num-candidates.txt b/vector-stores/output/05-elasticsearch-num-candidates.txt new file mode 100644 index 0000000..f3de5e1 --- /dev/null +++ b/vector-stores/output/05-elasticsearch-num-candidates.txt @@ -0,0 +1,7 @@ +Spring AI 2.0.1 sends num_candidates = (int)(1.5 * topK) = 15 for topK=10 (ElasticsearchVectorStore bytecode: ldc2_w 1.5d, dmul, d2i) +same index, same 100 queries, raw _search with the num_candidates varied: +num_candidates recall@10 p50 ms +15 0.491 4.5 +50 0.489 3.9 +100 0.490 3.8 +500 0.490 4.0 diff --git a/vector-stores/output/06-redis-ef-runtime.txt b/vector-stores/output/06-redis-ef-runtime.txt new file mode 100644 index 0000000..1d7a267 --- /dev/null +++ b/vector-stores/output/06-redis-ef-runtime.txt @@ -0,0 +1,5 @@ +Redis HNSW EF_RUNTIME: the library default (not set by Spring AI) against explicit values; index rebuilt and reloaded each time +ef_runtime recall@10 p50 ms filt recall filt p50 ms +default 0.616 0.6 1.000 1.6 +50 0.962 0.6 1.000 1.3 +200 1.000 0.8 1.000 1.3 diff --git a/vector-stores/output/07-pgvector-ef-search.txt b/vector-stores/output/07-pgvector-ef-search.txt new file mode 100644 index 0000000..3f60560 --- /dev/null +++ b/vector-stores/output/07-pgvector-ef-search.txt @@ -0,0 +1,7 @@ +pgvector 0.8.7: hnsw.ef_search against the filtered query (category == 'c3' && year >= 2022); the filter is applied AFTER the index scan +ef_search iterative_scan mean hits p50 ms filt recall +40 off 1.8 1.2 0.167 +200 off 2.0 1.7 0.198 +1000 off 10.0 28.2 1.000 +40 relaxed_order 10.0 7.7 0.611 +40 strict_order 10.0 9.7 0.281 diff --git a/vector-stores/output/08-qdrant-payload-index-and-threshold.txt b/vector-stores/output/08-qdrant-payload-index-and-threshold.txt new file mode 100644 index 0000000..70004c4 --- /dev/null +++ b/vector-stores/output/08-qdrant-payload-index-and-threshold.txt @@ -0,0 +1,4 @@ +filtered query, no payload index: recall 1.000 p50 26.2 ms p95 37.2 ms +filtered query, keyword+integer index: recall 1.000 p50 1.5 ms p95 3.2 ms +collection vs_small points_count=10000 segments_count=2 indexed_vectors_count=0 indexing_threshold=10000 KB +collection vs_bench points_count=30000 segments_count=2 indexed_vectors_count=30000 indexing_threshold=10000 KB diff --git a/vector-stores/output/09-redis-undeclared-metadata.txt b/vector-stores/output/09-redis-undeclared-metadata.txt new file mode 100644 index 0000000..a20b435 --- /dev/null +++ b/vector-stores/output/09-redis-undeclared-metadata.txt @@ -0,0 +1,2 @@ +metadata fields NOT declared; filter category == 'c3' threw IllegalArgumentException + Not allowed filter identifier name: category diff --git a/vector-stores/output/10-elasticsearch-index-options.txt b/vector-stores/output/10-elasticsearch-index-options.txt new file mode 100644 index 0000000..95495de --- /dev/null +++ b/vector-stores/output/10-elasticsearch-index-options.txt @@ -0,0 +1,6 @@ +same 30,000 documents, same queries; only the dense_vector index_options.type differs (the first row is Spring AI's own mapping) +type recall@10 p50 ms filt recall filt p50 ms store +bbq_hnsw* 0.491 3.6 0.544 3.0 54.4mb +int8_hnsw 0.952 3.0 0.974 3.1 64.1mb +hnsw 0.995 3.0 1.000 2.7 52.6mb +* default; its full mapping is in 04-defaults-that-matter.txt diff --git a/vector-stores/output/11-ops-footprint.txt b/vector-stores/output/11-ops-footprint.txt new file mode 100644 index 0000000..adca395 --- /dev/null +++ b/vector-stores/output/11-ops-footprint.txt @@ -0,0 +1,6 @@ +30000 x 384 float32 = 46 MB of raw vectors. PSS = proportional set size of the service's processes, read from /proc, after loading and querying. +store PSS MB size of the loaded data, as the store reports it +pgvector 153 table+toast 59 MB, hnsw index 59 MB +redis 422 used_memory 434 MB for 30000 keys (everything lives in RAM); one document = 10600 bytes; vector_index_sz_mb=55.4 +qdrant 172 storage dir 144 MB +elasticsearch 1550 index store.size/docs.count [{"store.size":"54.3mb","docs.count":"30000"}] (JVM heap fixed at -Xmx1g) diff --git a/vector-stores/output/12-redis-1x-compile.txt b/vector-stores/output/12-redis-1x-compile.txt new file mode 100644 index 0000000..5c48f37 --- /dev/null +++ b/vector-stores/output/12-redis-1x-compile.txt @@ -0,0 +1,13 @@ +$ javac -cp src/broken/Redis1xStyle.java +src/broken/Redis1xStyle.java:11: error: cannot find symbol + .vectorAlgorithm(RedisVectorStore.Algorithm.HSNW) + ^ + symbol: variable HSNW + location: class Algorithm +src/broken/Redis1xStyle.java:10: error: incompatible types: JedisPooled cannot be converted to RedisClient + return RedisVectorStore.builder(jedis, em) + ^ +Note: src/broken/Redis1xStyle.java uses or overrides a deprecated API. +Note: Recompile with -Xlint:deprecation for details. +Note: Some messages have been simplified; recompile with -Xdiags:verbose to get full output +2 errors diff --git a/vector-stores/output/13-es-disk-watermark.txt b/vector-stores/output/13-es-disk-watermark.txt new file mode 100644 index 0000000..1ab240d --- /dev/null +++ b/vector-stores/output/13-es-disk-watermark.txt @@ -0,0 +1,9 @@ +$ curl -s localhost:9200/_cat/shards/vs_es_int8_hnsw?h=index,shard,prirep,state,unassigned.reason +vs_es_int8_hnsw 0 p UNASSIGNED INDEX_CREATED +vs_es_int8_hnsw 0 r UNASSIGNED INDEX_CREATED + +$ curl -s localhost:9200/_cluster/allocation/explain (index, shard and decider lines) +index: vs_es_int8_hnsw | can_allocate: no | unassigned reason: INDEX_CREATED +disk_threshold -> NO | the node is above the high watermark cluster setting [cluster.routing.allocation.disk.watermark.high=90%], having less than the minimum required [25.1gb] free space, actual free: [25.1gb], actual used: [90%] +$ grep DiskThresholdMonitor eslogs/elasticsearch.log | tail -1 +[WARN ][o.e.c.r.a.DiskThresholdMonitor] [vm] high disk watermark [90%] exceeded on [UYHyLZrWSZSD9gDv_TWuGw][vm][/tmp/vs-run/esdata] free: 25.1gb[9.9%], shards will be relocated away from this node; currently relocating away shards totalling [0] bytes; the node is expected to continue to exceed the high disk watermark when these diff --git a/vector-stores/pom.xml b/vector-stores/pom.xml new file mode 100644 index 0000000..b7a9f27 --- /dev/null +++ b/vector-stores/pom.xml @@ -0,0 +1,65 @@ + + + 4.0.0 + + + org.springframework.boot + spring-boot-starter-parent + 4.1.1 + + + + com.ankurm + vector-stores + 1.0.0 + vector-stores + The same 30,000-document dataset behind Spring AI's VectorStore API on pgvector, Redis, Qdrant and Elasticsearch: ingest time, recall@10, latency, metadata filtering and what each costs to run. + + + 25 + 2.0.1 + + + + + + org.springframework.ai + spring-ai-bom + ${spring-ai.version} + pom + import + + + + + + + org.springframework.aispring-ai-pgvector-store + org.springframework.aispring-ai-redis-store + org.springframework.aispring-ai-qdrant-store + org.springframework.aispring-ai-elasticsearch-store + org.postgresqlpostgresql + org.springframework.bootspring-boot-starter-jdbc + + + org.springframework.boot + spring-boot-starter-test + test + + + + + + + org.apache.maven.plugins + maven-surefire-plugin + + -Duser.timezone=UTC -Dstdout.encoding=UTF-8 -Dfile.encoding=UTF-8 -Xmx3g + + + + + diff --git a/vector-stores/scripts/capture-1x-compile.sh b/vector-stores/scripts/capture-1x-compile.sh new file mode 100755 index 0000000..6141a41 --- /dev/null +++ b/vector-stores/scripts/capture-1x-compile.sh @@ -0,0 +1,12 @@ +#!/usr/bin/env bash +# Compiles src/broken/Redis1xStyle.java (the way the 1.x Redis store was wired) against Spring AI 2.0.1 +# and keeps the compiler's own message in output/12-redis-1x-compile.txt. +set -uo pipefail +cd "$(dirname "$0")/.." +mvn -q -B dependency:build-classpath -Dmdep.outputFile=target/classpath.txt >/dev/null 2>&1 +mkdir -p target/broken +{ + echo '$ javac -cp src/broken/Redis1xStyle.java' + javac -cp "$(cat target/classpath.txt)" -d target/broken src/broken/Redis1xStyle.java 2>&1 | grep -v -E '^Picked up' +} > output/12-redis-1x-compile.txt +cat output/12-redis-1x-compile.txt diff --git a/vector-stores/scripts/run-all.sh b/vector-stores/scripts/run-all.sh new file mode 100755 index 0000000..a54ea28 --- /dev/null +++ b/vector-stores/scripts/run-all.sh @@ -0,0 +1,9 @@ +#!/usr/bin/env bash +# Regenerates every file in output/. Needs the four services: scripts/services-up.sh (no Docker). +# A store whose port is closed is skipped, so the run still completes with fewer rows. +set -euo pipefail +cd "$(dirname "$0")/.." +mkdir -p output && find output -type f ! -name "13-*" -delete # 13 is a one-off capture, see its header +mvn -B -q clean test +scripts/capture-1x-compile.sh >/dev/null +ls output diff --git a/vector-stores/scripts/services-up.sh b/vector-stores/scripts/services-up.sh new file mode 100755 index 0000000..88a070e --- /dev/null +++ b/vector-stores/scripts/services-up.sh @@ -0,0 +1,54 @@ +#!/usr/bin/env bash +# Starts the four vector stores the benchmark compares, with no Docker: +# pgvector 127.0.0.1:5441 apt: postgresql-16 + postgresql-16-pgvector (runs as the postgres user) +# Redis Stack 127.0.0.1:6392 tarball: redis-stack-server (Redis + RediSearch module) +# Qdrant 127.0.0.1:6334 tarball: single static binary (gRPC 6334, REST 6333) +# Elasticsearch 127.0.0.1:9200 tarball: single node, security off (runs as the 'es' user) +# TOOLS defaults to /tmp/tools; see the article for the exact download URLs. +set -euo pipefail +TOOLS="${TOOLS:-/tmp/tools}" +RUN=/tmp/vs-run; mkdir -p "$RUN"; chmod 777 "$RUN" +listening() { (exec 3<>/dev/tcp/127.0.0.1/$1) 2>/dev/null; } + +# ---- pgvector ------------------------------------------------------------------------------- +PGBIN=/usr/lib/postgresql/16/bin; PGDATA=$RUN/pgdata; PGPORT=5441 +if ! listening $PGPORT; then + id postgres >/dev/null 2>&1 + rm -rf "$PGDATA"; mkdir -p "$PGDATA"; chown postgres "$PGDATA" + su postgres -c "$PGBIN/initdb -D $PGDATA -A trust >/dev/null" + su postgres -c "$PGBIN/pg_ctl -D $PGDATA -o '-p $PGPORT -c listen_addresses=127.0.0.1 -c unix_socket_directories=$RUN' -l $RUN/pg.log -w start >/dev/null" + su postgres -c "psql -h $RUN -p $PGPORT -d postgres -c \"CREATE ROLE vs LOGIN PASSWORD 'vs' SUPERUSER\" -c 'CREATE DATABASE vsdb OWNER vs'" >/dev/null +fi + +# ---- Redis Stack ---------------------------------------------------------------------------- +RS="$TOOLS/redis-stack-server-7.4.0-v8" +if ! listening 6392; then + mkdir -p "$RUN/redis" + "$RS/bin/redis-server" --port 6392 --bind 127.0.0.1 --dir "$RUN/redis" --save "" --daemonize yes \ + --loadmodule "$RS/lib/redisearch.so" --loadmodule "$RS/lib/rejson.so" --logfile "$RUN/redis.log" >/dev/null +fi + +# ---- Qdrant --------------------------------------------------------------------------------- +if ! listening 6334; then + mkdir -p "$RUN/qdrant" + ( cd "$RUN/qdrant" && QDRANT__STORAGE__STORAGE_PATH="$RUN/qdrant/storage" QDRANT__SERVICE__HOST=127.0.0.1 \ + setsid nohup "$TOOLS/qdrant/qdrant" > "$RUN/qdrant.log" 2>&1 < /dev/null & ) +fi + +# ---- Elasticsearch -------------------------------------------------------------------------- +ES="$TOOLS/$(ls "$TOOLS" | grep -m1 '^elasticsearch-')" +if ! listening 9200; then + id es >/dev/null 2>&1 || useradd -m es + chown -R es "$ES" + su es -c "ES_JAVA_OPTS='-Xms1g -Xmx1g' $ES/bin/elasticsearch -d -p $RUN/es.pid \ + -E xpack.security.enabled=false -E discovery.type=single-node -E http.host=127.0.0.1 \ + -E path.data=$RUN/esdata -E path.logs=$RUN/eslogs" >/dev/null 2>&1 || true +fi + +for i in $(seq 1 90); do + curl -s -o /dev/null 127.0.0.1:9200 && curl -s -o /dev/null 127.0.0.1:6333 && break; sleep 2 +done +# On a disk that is >90% full Elasticsearch silently refuses to allocate shards (see output/13-es-disk-watermark.txt). +curl -s -X PUT -H 'Content-Type: application/json' 127.0.0.1:9200/_cluster/settings \ + -d '{"persistent":{"cluster.routing.allocation.disk.threshold_enabled":false}}' >/dev/null || true +echo "pgvector :$PGPORT (vsdb/vs/vs), redis-stack :6392, qdrant :6334 (REST :6333), elasticsearch :9200" diff --git a/vector-stores/src/broken/Redis1xStyle.java b/vector-stores/src/broken/Redis1xStyle.java new file mode 100644 index 0000000..b8cbc62 --- /dev/null +++ b/vector-stores/src/broken/Redis1xStyle.java @@ -0,0 +1,14 @@ +// Not compiled by the build (it lives outside src/main and src/test). scripts/capture-1x-compile.sh +// compiles it against Spring AI 2.0.1 and captures the compiler's own message. +import org.springframework.ai.embedding.EmbeddingModel; +import org.springframework.ai.vectorstore.redis.RedisVectorStore; +import redis.clients.jedis.JedisPooled; + +class Redis1xStyle { + RedisVectorStore build(EmbeddingModel em) { + JedisPooled jedis = new JedisPooled("127.0.0.1", 6392); + return RedisVectorStore.builder(jedis, em) + .vectorAlgorithm(RedisVectorStore.Algorithm.HSNW) + .build(); + } +} diff --git a/vector-stores/src/main/java/com/ankurm/vectorstores/Dataset.java b/vector-stores/src/main/java/com/ankurm/vectorstores/Dataset.java new file mode 100644 index 0000000..23a7a4f --- /dev/null +++ b/vector-stores/src/main/java/com/ankurm/vectorstores/Dataset.java @@ -0,0 +1,120 @@ +package com.ankurm.vectorstores; + +import java.util.ArrayList; +import java.util.Arrays; +import java.util.LinkedHashMap; +import java.util.List; +import java.util.Map; +import java.util.Random; +import java.util.UUID; + +import org.springframework.ai.document.Document; + +/** + * A synthetic, clustered, unit-length 384-dimension dataset with three metadata fields. + * + *

Synthetic on purpose: with no API key and no model download, a hashing "embedding" would + * have no structure and every ANN index would look the same. Clusters give HNSW real + * neighbourhoods to find. What this measures is the store (index, filter, wire), not + * how good any embedding model is. + */ +public final class Dataset { + + public static final int DIMS = 384; + public static final int CLUSTERS = 40; + + public final int size; + public final float[][] docVectors; + public final float[][] queryVectors; + public final String[] category; // "c0".."c7", correlated with the cluster + public final int[] year; // 2015..2025 + public final String[] region; // "in","eu","us","ap" + private final Map idToIndex = new LinkedHashMap<>(); + + public Dataset(int size, int queries, long seed) { + this.size = size; + Random r = new Random(seed); + float[][] centroids = new float[CLUSTERS][DIMS]; + for (float[] c : centroids) { + for (int d = 0; d < DIMS; d++) c[d] = (float) r.nextGaussian(); + normalize(c); + } + docVectors = new float[size][]; + category = new String[size]; + year = new int[size]; + region = new String[size]; + String[] regions = {"in", "eu", "us", "ap"}; + for (int i = 0; i < size; i++) { + int cl = i % CLUSTERS; + docVectors[i] = noisy(centroids[cl], r, 0.30f); + category[i] = "c" + (cl % 8); + year[i] = 2015 + r.nextInt(11); + region[i] = regions[r.nextInt(4)]; + idToIndex.put(idOf(i), i); + } + queryVectors = new float[queries][]; + for (int q = 0; q < queries; q++) queryVectors[q] = noisy(centroids[r.nextInt(CLUSTERS)], r, 0.30f); + } + + private static float[] noisy(float[] centroid, Random r, float sigma) { + float[] v = new float[DIMS]; + for (int d = 0; d < DIMS; d++) v[d] = centroid[d] + sigma * (float) r.nextGaussian() / (float) Math.sqrt(DIMS) * 4f; + normalize(v); + return v; + } + + private static void normalize(float[] v) { + double s = 0; + for (float x : v) s += x * x; + float n = (float) Math.sqrt(s); + for (int i = 0; i < v.length; i++) v[i] /= n; + } + + public static UUID idOf(int i) { + return UUID.nameUUIDFromBytes(("doc-" + i).getBytes()); + } + + public int indexOf(String id) { + Integer i = idToIndex.get(UUID.fromString(id)); + return i == null ? -1 : i; + } + + /** The vector the {@link LookupEmbeddingModel} returns for "d123" (a document) or "q7" (a query). */ + public float[] vectorForText(String text) { + int n = Integer.parseInt(text.substring(1)); + return text.charAt(0) == 'd' ? docVectors[n] : queryVectors[n]; + } + + public List documents() { + List docs = new ArrayList<>(size); + for (int i = 0; i < size; i++) { + docs.add(Document.builder() + .id(idOf(i).toString()) + .text("d" + i) + .metadata(Map.of("category", category[i], "year", year[i], "region", region[i])) + .build()); + } + return docs; + } + + /** Exact top-k by cosine (vectors are unit length, so dot product), optionally restricted. */ + public int[] exactTopK(int query, int k, java.util.function.IntPredicate allowed) { + float[] q = queryVectors[query]; + double[] best = new double[k]; + int[] bestIdx = new int[k]; + Arrays.fill(best, -2); + Arrays.fill(bestIdx, -1); + for (int i = 0; i < size; i++) { + if (allowed != null && !allowed.test(i)) continue; + double s = 0; + float[] v = docVectors[i]; + for (int d = 0; d < DIMS; d++) s += q[d] * v[d]; + if (s > best[k - 1]) { + int p = k - 1; + while (p > 0 && best[p - 1] < s) { best[p] = best[p - 1]; bestIdx[p] = bestIdx[p - 1]; p--; } + best[p] = s; bestIdx[p] = i; + } + } + return bestIdx; + } +} diff --git a/vector-stores/src/main/java/com/ankurm/vectorstores/LookupEmbeddingModel.java b/vector-stores/src/main/java/com/ankurm/vectorstores/LookupEmbeddingModel.java new file mode 100644 index 0000000..e0fbd6e --- /dev/null +++ b/vector-stores/src/main/java/com/ankurm/vectorstores/LookupEmbeddingModel.java @@ -0,0 +1,46 @@ +package com.ankurm.vectorstores; + +import java.util.ArrayList; +import java.util.List; + +import org.springframework.ai.document.Document; +import org.springframework.ai.embedding.Embedding; +import org.springframework.ai.embedding.EmbeddingModel; +import org.springframework.ai.embedding.EmbeddingRequest; +import org.springframework.ai.embedding.EmbeddingResponse; + +/** + * An {@link EmbeddingModel} that looks vectors up in a {@link Dataset} ("d12" is document 12, + * "q3" is query 3). It is how all four stores get the same vectors with no network call, + * and it counts calls so a test can show how often each store re-embeds. + */ +public final class LookupEmbeddingModel implements EmbeddingModel { + + private final Dataset dataset; + public int calls; + + public LookupEmbeddingModel(Dataset dataset) { + this.dataset = dataset; + } + + @Override + public EmbeddingResponse call(EmbeddingRequest request) { + List out = new ArrayList<>(); + int i = 0; + for (String text : request.getInstructions()) { + out.add(new Embedding(dataset.vectorForText(text), i++)); + } + calls++; + return new EmbeddingResponse(out); + } + + @Override + public float[] embed(Document document) { + return dataset.vectorForText(document.getText()); + } + + @Override + public int dimensions() { + return Dataset.DIMS; + } +} diff --git a/vector-stores/src/main/java/com/ankurm/vectorstores/StoreFactory.java b/vector-stores/src/main/java/com/ankurm/vectorstores/StoreFactory.java new file mode 100644 index 0000000..763039b --- /dev/null +++ b/vector-stores/src/main/java/com/ankurm/vectorstores/StoreFactory.java @@ -0,0 +1,137 @@ +package com.ankurm.vectorstores; + +import java.net.URI; +import java.net.http.HttpClient; +import java.net.http.HttpRequest; +import java.net.http.HttpResponse; + +import co.elastic.clients.transport.rest5_client.low_level.Rest5Client; +import io.qdrant.client.QdrantClient; +import io.qdrant.client.QdrantGrpcClient; +import org.apache.hc.core5.http.HttpHost; +import org.springframework.ai.embedding.EmbeddingModel; +import org.springframework.ai.vectorstore.VectorStore; +import org.springframework.ai.vectorstore.elasticsearch.ElasticsearchVectorStore; +import org.springframework.ai.vectorstore.elasticsearch.ElasticsearchVectorStoreOptions; +import org.springframework.ai.vectorstore.pgvector.PgVectorStore; +import org.springframework.ai.vectorstore.qdrant.QdrantVectorStore; +import org.springframework.ai.vectorstore.redis.RedisVectorStore; +import org.springframework.jdbc.core.JdbcTemplate; +import org.springframework.jdbc.datasource.DriverManagerDataSource; +import redis.clients.jedis.RedisClient; + +/** + * Builds the four {@link VectorStore}s with the library's own defaults wherever possible, so the + * comparison is "what you get if you follow the reference page", not a tuned showcase. The few + * explicit settings (dimension, cosine, the Redis metadata schema) are the ones the stores require. + */ +public final class StoreFactory { + + public static final String PG = "pgvector", REDIS = "redis", QDRANT = "qdrant", ES = "elasticsearch"; + public static final String NAME = "vs_bench"; + + private StoreFactory() { } + + public static boolean up(int port) { + try (var s = new java.net.Socket("127.0.0.1", port)) { return true; } catch (Exception e) { return false; } + } + + public static boolean available(String store) { + return switch (store) { + case PG -> up(5441); + case REDIS -> up(6392); + case QDRANT -> up(6334) && up(6333); + case ES -> up(9200); + default -> false; + }; + } + + public static JdbcTemplate pgJdbc() { + var ds = new DriverManagerDataSource("jdbc:postgresql://127.0.0.1:5441/vsdb", "vs", "vs"); + return new JdbcTemplate(ds); + } + + public static VectorStore pgvector(EmbeddingModel em) { + JdbcTemplate jdbc = pgJdbc(); + jdbc.execute("CREATE EXTENSION IF NOT EXISTS vector"); + return init(PgVectorStore.builder(jdbc, em) + .vectorTableName(NAME) + .dimensions(Dataset.DIMS) + .distanceType(PgVectorStore.PgDistanceType.COSINE_DISTANCE) + .indexType(PgVectorStore.PgIndexType.HNSW) + .removeExistingVectorStoreTable(true) + .initializeSchema(true) + .build()); + } + + public static RedisClient jedis() { + return RedisClient.create("127.0.0.1", 6392); + } + + /** {@code efRuntime == null} means the library default. */ + public static VectorStore redis(RedisClient jedis, EmbeddingModel em, Integer efRuntime, boolean declareMetadata) { + try { jedis.ftDropIndex(NAME); } catch (Exception ignored) { } + jedis.flushAll(); + var b = RedisVectorStore.builder(jedis, em) + .indexName(NAME) + .prefix(NAME + ":") + .vectorAlgorithm(RedisVectorStore.Algorithm.HNSW) + .distanceMetric(RedisVectorStore.DistanceMetric.COSINE) + .initializeSchema(true); + if (declareMetadata) { + b.metadataFields(RedisVectorStore.MetadataField.tag("category"), + RedisVectorStore.MetadataField.numeric("year"), + RedisVectorStore.MetadataField.tag("region")); + } + if (efRuntime != null) b.hnswEfRuntime(efRuntime); + return init(b.build()); + } + + public static QdrantClient qdrantClient() { + return new QdrantClient(QdrantGrpcClient.newBuilder("127.0.0.1", 6334, false).build()); + } + + public static VectorStore qdrant(QdrantClient client, EmbeddingModel em) throws Exception { + http("DELETE", "http://127.0.0.1:6333/collections/" + NAME, null); + return init(QdrantVectorStore.builder(client, em) + .collectionName(NAME) + .initializeSchema(true) + .build()); + } + + public static Rest5Client esClient() throws Exception { + return Rest5Client.builder(HttpHost.create("http://127.0.0.1:9200")).build(); + } + + public static VectorStore elasticsearch(Rest5Client client, EmbeddingModel em) throws Exception { + http("DELETE", "http://127.0.0.1:9200/" + NAME, null); + var o = new ElasticsearchVectorStoreOptions(); + o.setIndexName(NAME); + o.setDimensions(Dataset.DIMS); + o.setSimilarity(org.springframework.ai.vectorstore.elasticsearch.SimilarityFunction.cosine); + VectorStore vs = init(ElasticsearchVectorStore.builder(client, em).options(o).initializeSchema(true).build()); + // A just-created index may not have its primary shard yet; a bulk sent too early fails after 1 minute. + http("GET", "http://127.0.0.1:9200/_cluster/health/" + NAME + "?wait_for_status=yellow&timeout=60s", null); + return vs; + } + + /** + * Built by hand, a store has not created its schema yet: the table, index or collection is made in + * {@code afterPropertiesSet()}, which Spring calls for a bean and nobody calls for a {@code build()}. + */ + static VectorStore init(VectorStore store) { + try { + ((org.springframework.beans.factory.InitializingBean) store).afterPropertiesSet(); + } catch (Exception e) { + throw new IllegalStateException(e); + } + return store; + } + + public static String http(String method, String url, String body) throws Exception { + var req = HttpRequest.newBuilder(URI.create(url)).header("Content-Type", "application/json") + .method(method, body == null ? HttpRequest.BodyPublishers.noBody() : HttpRequest.BodyPublishers.ofString(body)) + .build(); + return HttpClient.newHttpClient().send(req, HttpResponse.BodyHandlers.ofString()).body(); + } +} diff --git a/vector-stores/src/test/java/com/ankurm/vectorstores/StoreComparisonTest.java b/vector-stores/src/test/java/com/ankurm/vectorstores/StoreComparisonTest.java new file mode 100644 index 0000000..77b799f --- /dev/null +++ b/vector-stores/src/test/java/com/ankurm/vectorstores/StoreComparisonTest.java @@ -0,0 +1,431 @@ +package com.ankurm.vectorstores; + +import java.util.ArrayList; +import java.util.Arrays; +import java.util.HashSet; +import java.util.LinkedHashMap; +import java.util.List; +import java.util.Map; +import java.util.Set; + +import co.elastic.clients.transport.rest5_client.low_level.Rest5Client; +import com.ankurm.vectorstores.support.Transcript; +import io.qdrant.client.QdrantClient; +import org.junit.jupiter.api.BeforeAll; +import org.junit.jupiter.api.MethodOrderer; +import org.junit.jupiter.api.Order; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.TestInstance; +import org.junit.jupiter.api.TestMethodOrder; +import org.springframework.ai.document.Document; +import org.springframework.ai.vectorstore.SearchRequest; +import org.springframework.ai.vectorstore.VectorStore; +import redis.clients.jedis.RedisClient; + +import static org.assertj.core.api.Assertions.assertThat; +import static org.junit.jupiter.api.Assumptions.assumeTrue; + +/** + * One 30,000-document dataset, four stores, library defaults. Each test writes a transcript to + * output/ and asserts the same numbers it prints, so the article's tables are assertions. + * A store whose service is not listening is skipped, not failed. + */ +@TestInstance(TestInstance.Lifecycle.PER_CLASS) +@TestMethodOrder(MethodOrderer.OrderAnnotation.class) +class StoreComparisonTest { + + static final int N = 30_000, QUERIES = 100, WARMUP = 20, K = 10, BATCH = 1_000; + static final String FILTER = "category == 'c3' && year >= 2022"; + static final List ALL = List.of(StoreFactory.PG, StoreFactory.REDIS, StoreFactory.QDRANT, StoreFactory.ES); + + Dataset ds; + LookupEmbeddingModel em; + final Map stores = new LinkedHashMap<>(); + final Map ingestSeconds = new LinkedHashMap<>(); + RedisClient jedis; + QdrantClient qdrant; + Rest5Client es; + + @BeforeAll + void loadEverything() throws Exception { + ds = new Dataset(N, QUERIES + WARMUP, 42L); + em = new LookupEmbeddingModel(ds); + jedis = StoreFactory.jedis(); + qdrant = StoreFactory.qdrantClient(); + es = StoreFactory.esClient(); + for (String s : ALL) { + if (!StoreFactory.available(s)) continue; + VectorStore vs = switch (s) { + case StoreFactory.PG -> StoreFactory.pgvector(em); + case StoreFactory.REDIS -> StoreFactory.redis(jedis, em, null, true); + case StoreFactory.QDRANT -> StoreFactory.qdrant(qdrant, em); + default -> StoreFactory.elasticsearch(es, em); + }; + stores.put(s, vs); + ingestSeconds.put(s, ingest(vs, ds.documents())); + awaitSearchable(s); + } + } + + double ingest(VectorStore vs, List docs) { + long t0 = System.nanoTime(); + for (int i = 0; i < docs.size(); i += BATCH) { + vs.add(docs.subList(i, Math.min(i + BATCH, docs.size()))); + } + return (System.nanoTime() - t0) / 1e9; + } + + /** Ingest returning is not the same as "searchable": ES needs a refresh, Qdrant an optimizer pass, Redis an index drain. */ + void awaitSearchable(String store) throws Exception { + switch (store) { + case StoreFactory.ES -> StoreFactory.http("POST", "http://127.0.0.1:9200/" + StoreFactory.NAME + "/_refresh", null); + case StoreFactory.QDRANT -> { + for (int i = 0; i < 120; i++) { + String info = StoreFactory.http("GET", "http://127.0.0.1:6333/collections/" + StoreFactory.NAME, null); + if (info.contains("\"status\":\"green\"") && info.contains("\"optimizer_status\":\"ok\"")) break; + Thread.sleep(500); + } + } + case StoreFactory.REDIS -> { + for (int i = 0; i < 120; i++) { + Object pct = jedis.ftInfo(StoreFactory.NAME).get("percent_indexed"); + if (pct != null && Double.parseDouble(pct.toString()) >= 1.0) break; + Thread.sleep(500); + } + } + default -> { } + } + } + + record Result(double recall, double p50, double p95, double meanHits) { } + + Result measure(VectorStore vs, String filter, boolean filtered) { + for (int q = 0; q < WARMUP; q++) vs.similaritySearch(request(q, filter)); + List lat = new ArrayList<>(); + double recall = 0, hits = 0; + for (int q = WARMUP; q < WARMUP + QUERIES; q++) { + long t0 = System.nanoTime(); + List res = vs.similaritySearch(request(q, filter)); + lat.add((System.nanoTime() - t0) / 1e6); + int[] truth = ds.exactTopK(q, K, filtered ? i -> ds.category[i].equals("c3") && ds.year[i] >= 2022 : null); + Set want = new HashSet<>(); + for (int t : truth) if (t >= 0) want.add(t); + int found = 0; + for (Document d : res) if (want.contains(ds.indexOf(d.getId()))) found++; + recall += want.isEmpty() ? 1.0 : (double) found / want.size(); + hits += res.size(); + } + lat.sort(Double::compare); + return new Result(recall / QUERIES, lat.get(QUERIES / 2), lat.get((int) (QUERIES * 0.95) - 1), hits / QUERIES); + } + + SearchRequest request(int q, String filter) { + var b = SearchRequest.builder().query("q" + q).topK(K).similarityThreshold(0.0); + if (filter != null) b.filterExpression(filter); + return b.build(); + } + + @Test @Order(1) + void ingestTimes() { + var t = new Transcript("01-ingest.txt"); + t.line("dataset: %d documents x %d dims, %d clusters, batches of %d; embeddings are looked up, so this is store time only", N, Dataset.DIMS, Dataset.CLUSTERS, BATCH); + t.line("%-14s %10s %12s", "store", "seconds", "docs/second"); + for (var e : ingestSeconds.entrySet()) { + t.line("%-14s %10.1f %12.0f", e.getKey(), e.getValue(), N / e.getValue()); + assertThat(e.getValue()).isPositive(); + } + t.write(); + assumeTrue(!stores.isEmpty()); + } + + @Test @Order(2) + void unfilteredRecallAndLatency() { + var t = new Transcript("02-unfiltered.txt"); + t.line("%d queries after %d warm-up, topK=%d, cosine, each store's library defaults", QUERIES, WARMUP, K); + t.line("%-14s %10s %9s %9s %10s", "store", "recall@10", "p50 ms", "p95 ms", "mean hits"); + for (var e : stores.entrySet()) { + Result r = measure(e.getValue(), null, false); + t.line("%-14s %10.3f %9.1f %9.1f %10.1f", e.getKey(), r.recall, r.p50, r.p95, r.meanHits); + assertThat(r.recall).isBetween(0.0, 1.0); + } + t.write(); + } + + @Test @Order(3) + void filteredRecallAndLatency() { + var t = new Transcript("03-filtered.txt"); + long matching = java.util.stream.IntStream.range(0, N).filter(i -> ds.category[i].equals("c3") && ds.year[i] >= 2022).count(); + t.line("filter: %s (%d of %d documents match, %.1f%%)", FILTER, matching, N, 100.0 * matching / N); + t.line("%-14s %10s %9s %9s %10s", "store", "recall@10", "p50 ms", "p95 ms", "mean hits"); + for (var e : stores.entrySet()) { + Result r = measure(e.getValue(), FILTER, true); + t.line("%-14s %10.3f %9.1f %9.1f %10.1f", e.getKey(), r.recall, r.p50, r.p95, r.meanHits); + assertThat(r.meanHits).isLessThanOrEqualTo(K); + } + t.write(); + } + + @Test @Order(4) + void settingsThatSurprise() throws Exception { + var t = new Transcript("04-defaults-that-matter.txt"); + if (StoreFactory.available(StoreFactory.PG)) { + String ef = StoreFactory.pgJdbc().execute((java.sql.Connection c) -> { + try (var st = c.createStatement()) { + st.execute("SELECT '[1,2,3]'::vector"); // loads the extension library, which registers hnsw.ef_search + var rs = st.executeQuery("SHOW hnsw.ef_search"); + rs.next(); + return rs.getString(1); + } + }); + t.line("pgvector hnsw.ef_search = %s", ef); + t.line("pgvector index: %s", StoreFactory.pgJdbc().queryForObject( + "SELECT indexdef FROM pg_indexes WHERE tablename='vs_bench' AND indexdef LIKE '%hnsw%'", String.class)); + } + if (StoreFactory.available(StoreFactory.ES)) { + String mapping = StoreFactory.http("GET", "http://127.0.0.1:9200/vs_bench/_mapping", null); + t.line("elasticsearch embedding mapping: %s", object(mapping, "embedding")); + } + if (StoreFactory.available(StoreFactory.QDRANT)) { + String info = StoreFactory.http("GET", "http://127.0.0.1:6333/collections/vs_bench", null); + t.line("qdrant points_count=%s indexed_vectors_count=%s", field(info, "points_count"), field(info, "indexed_vectors_count")); + t.line("qdrant hnsw_config: %s", info.replaceAll(".*(\"hnsw_config\":\\{[^}]*\\}).*", "$1")); + t.line("qdrant indexing_threshold (KB): %s", field(info, "indexing_threshold")); + } + if (StoreFactory.available(StoreFactory.REDIS)) { + Map info = jedis.ftInfo(StoreFactory.NAME); + t.line("redis num_docs=%s percent_indexed=%s", info.get("num_docs"), info.get("percent_indexed")); + } + t.write(); + } + + + // ---- what it costs to keep running ------------------------------------------------------- + + @Test @Order(4) + void operationalFootprint() throws Exception { + var t = new Transcript("11-ops-footprint.txt"); + long rawMb = (long) N * Dataset.DIMS * 4 / 1_000_000; + t.line("%d x %d float32 = %d MB of raw vectors. PSS = proportional set size of the service's processes, read from /proc, after loading and querying.", N, Dataset.DIMS, rawMb); + t.line("%-14s %9s %s", "store", "PSS MB", "size of the loaded data, as the store reports it"); + if (stores.containsKey(StoreFactory.PG)) { + var j = StoreFactory.pgJdbc(); + t.line("%-14s %9d table+toast %s, hnsw index %s", "pgvector", pss("postgres"), + j.queryForObject("SELECT pg_size_pretty(pg_table_size('vs_bench'))", String.class), + j.queryForObject("SELECT pg_size_pretty(pg_relation_size('vs_bench_index'))", String.class)); + } + if (stores.containsKey(StoreFactory.REDIS)) { + Map info = jedis.ftInfo(StoreFactory.NAME); + String firstKey = jedis.scan("0", new redis.clients.jedis.params.ScanParams().match(StoreFactory.NAME + ":*").count(10)).getResult().get(0); + t.line("%-14s %9d used_memory %d MB for %d keys (everything lives in RAM); one document = %d bytes; vector_index_sz_mb=%.1f", "redis", pss("redis-server"), + Long.parseLong(redisInfo("used_memory")) / 1_000_000, jedis.dbSize(), jedis.memoryUsage(firstKey), Double.parseDouble(info.get("vector_index_sz_mb").toString())); + } + if (stores.containsKey(StoreFactory.QDRANT)) { + t.line("%-14s %9d storage dir %s MB", "qdrant", pss("tools/qdrant/qdrant"), shell("du -sm /tmp/vs-run/qdrant/storage | cut -f1")); + } + if (stores.containsKey(StoreFactory.ES)) { + String cat = StoreFactory.http("GET", "http://127.0.0.1:9200/_cat/indices/vs_bench?h=store.size,docs.count&format=json", null); + t.line("%-14s %9d index store.size/docs.count %s (JVM heap fixed at -Xmx1g)", "elasticsearch", pss("org.elasticsearch.server"), cat); + } + t.write(); + } + + static long pss(String commandLineFragment) { + return ProcessHandle.allProcesses() + .filter(p -> p.info().commandLine().map(c -> c.contains(commandLineFragment)).orElse(false) + && (!commandLineFragment.equals("postgres") || p.info().user().map("postgres"::equals).orElse(false))) + .mapToLong(p -> { + try { + for (String l : java.nio.file.Files.readAllLines(java.nio.file.Path.of("/proc/" + p.pid() + "/smaps_rollup"))) { + if (l.startsWith("Pss:")) return Long.parseLong(l.replaceAll("\\D+", "")) / 1024; + } + } catch (Exception ignored) { } + return 0; + }).sum(); + } + + String redisInfo(String key) { + return jedis.info("memory").lines().filter(l -> l.startsWith(key + ":")).map(l -> l.substring(key.length() + 1).trim()).findFirst().orElse("0"); + } + + static String shell(String cmd) throws Exception { + Process pr = new ProcessBuilder("bash", "-c", cmd).redirectErrorStream(true).start(); + return new String(pr.getInputStream().readAllBytes()).trim(); + } + + + // ---- why the numbers look the way they do ------------------------------------------------ + + @Test @Order(5) + void elasticsearchCandidatesNotQuantisation() throws Exception { + assumeTrue(stores.containsKey(StoreFactory.ES)); + var t = new Transcript("05-elasticsearch-num-candidates.txt"); + t.line("Spring AI 2.0.1 sends num_candidates = (int)(1.5 * topK) = %d for topK=%d (ElasticsearchVectorStore bytecode: ldc2_w 1.5d, dmul, d2i)", (int) (1.5 * K), K); + t.line("same index, same 100 queries, raw _search with the num_candidates varied:"); + t.line("%-16s %10s %9s", "num_candidates", "recall@10", "p50 ms"); + for (int nc : new int[] {15, 50, 100, 500}) { + double recall = 0; + List lat = new ArrayList<>(); + for (int q = WARMUP; q < WARMUP + QUERIES; q++) { + String body = "{\"size\":10,\"_source\":false,\"knn\":{\"field\":\"embedding\",\"k\":10,\"num_candidates\":" + nc + + ",\"query_vector\":" + Arrays.toString(ds.queryVectors[q]) + "}}"; + long t0 = System.nanoTime(); + String res = StoreFactory.http("POST", "http://127.0.0.1:9200/vs_bench/_search", body); + lat.add((System.nanoTime() - t0) / 1e6); + Set want = new HashSet<>(); + for (int x : ds.exactTopK(q, K, null)) want.add(x); + var m = java.util.regex.Pattern.compile("\"_id\":\"([^\"]+)\"").matcher(res); + int found = 0; + while (m.find()) if (want.contains(ds.indexOf(m.group(1)))) found++; + recall += (double) found / K; + } + lat.sort(Double::compare); + t.line("%-16d %10.3f %9.1f", nc, recall / QUERIES, lat.get(QUERIES / 2)); + } + t.write(); + } + + @Test @Order(6) + void redisEfRuntime() throws Exception { + assumeTrue(stores.containsKey(StoreFactory.REDIS)); + var t = new Transcript("06-redis-ef-runtime.txt"); + t.line("Redis HNSW EF_RUNTIME: the library default (not set by Spring AI) against explicit values; index rebuilt and reloaded each time"); + t.line("%-12s %10s %9s %12s %14s", "ef_runtime", "recall@10", "p50 ms", "filt recall", "filt p50 ms"); + for (Integer ef : new Integer[] {null, 50, 200}) { + VectorStore vs = StoreFactory.redis(jedis, em, ef, true); + ingest(vs, ds.documents()); + awaitSearchable(StoreFactory.REDIS); + Result u = measure(vs, null, false); + Result f = measure(vs, FILTER, true); + t.line("%-12s %10.3f %9.1f %12.3f %14.1f", ef == null ? "default" : ef, u.recall, u.p50, f.recall, f.p50); + } + t.write(); + } + + @Test @Order(7) + void pgvectorEfSearchAndFilters() throws Exception { + assumeTrue(stores.containsKey(StoreFactory.PG)); + var t = new Transcript("07-pgvector-ef-search.txt"); + t.line("pgvector %s: hnsw.ef_search against the filtered query (%s); the filter is applied AFTER the index scan", + StoreFactory.pgJdbc().queryForObject("SELECT extversion FROM pg_extension WHERE extname='vector'", String.class), FILTER); + t.line("%-10s %-14s %10s %9s %12s", "ef_search", "iterative_scan", "mean hits", "p50 ms", "filt recall"); + Object[][] variants = {{40, "off"}, {200, "off"}, {1000, "off"}, {40, "relaxed_order"}, {40, "strict_order"}}; + for (Object[] v : variants) { + var one = new org.springframework.jdbc.datasource.SingleConnectionDataSource("jdbc:postgresql://127.0.0.1:5441/vsdb", "vs", "vs", true); + var jdbc = new org.springframework.jdbc.core.JdbcTemplate(one); + jdbc.execute("SELECT '[1,2,3]'::vector"); + jdbc.execute("SET hnsw.ef_search = " + v[0]); + jdbc.execute("SET hnsw.iterative_scan = " + v[1]); + VectorStore vs = org.springframework.ai.vectorstore.pgvector.PgVectorStore.builder(jdbc, em) + .vectorTableName(StoreFactory.NAME).dimensions(Dataset.DIMS) + .distanceType(org.springframework.ai.vectorstore.pgvector.PgVectorStore.PgDistanceType.COSINE_DISTANCE) + .initializeSchema(false).vectorTableValidationsEnabled(false).build(); + Result f = measure(vs, FILTER, true); + t.line("%-10d %-14s %10.1f %9.1f %12.3f", (Integer) v[0], v[1], f.meanHits, f.p50, f.recall); + one.destroy(); + } + t.write(); + } + + @Test @Order(8) + void qdrantPayloadIndexAndIndexingThreshold() throws Exception { + assumeTrue(stores.containsKey(StoreFactory.QDRANT)); + var t = new Transcript("08-qdrant-payload-index-and-threshold.txt"); + VectorStore vs = stores.get(StoreFactory.QDRANT); + Result before = measure(vs, FILTER, true); + t.line("filtered query, no payload index: recall %.3f p50 %.1f ms p95 %.1f ms", before.recall, before.p50, before.p95); + String base = "http://127.0.0.1:6333/collections/" + StoreFactory.NAME + "/index?wait=true"; + StoreFactory.http("PUT", base, "{\"field_name\":\"category\",\"field_schema\":\"keyword\"}"); + StoreFactory.http("PUT", base, "{\"field_name\":\"year\",\"field_schema\":\"integer\"}"); + awaitSearchable(StoreFactory.QDRANT); + Result after = measure(vs, FILTER, true); + t.line("filtered query, keyword+integer index: recall %.3f p50 %.1f ms p95 %.1f ms", after.recall, after.p50, after.p95); + + // The indexing threshold: below ~20,000 KB of vectors Qdrant never builds HNSW at all. + var small = new LookupEmbeddingModel(ds); + VectorStore tiny = org.springframework.ai.vectorstore.qdrant.QdrantVectorStore.builder(qdrant, small) + .collectionName("vs_small").initializeSchema(true).build(); + ((org.springframework.beans.factory.InitializingBean) tiny).afterPropertiesSet(); + StoreFactory.http("DELETE", "http://127.0.0.1:6333/collections/vs_small", null); + tiny = org.springframework.ai.vectorstore.qdrant.QdrantVectorStore.builder(qdrant, small) + .collectionName("vs_small").initializeSchema(true).build(); + StoreFactory.init(tiny); + ingest(tiny, ds.documents().subList(0, 10_000)); + Thread.sleep(15_000); + for (String c : new String[] {"vs_small", StoreFactory.NAME}) { + String info = StoreFactory.http("GET", "http://127.0.0.1:6333/collections/" + c, null); + t.line("collection %-9s points_count=%-6s segments_count=%-3s indexed_vectors_count=%-6s indexing_threshold=%s KB", + c, field(info, "points_count"), field(info, "segments_count"), field(info, "indexed_vectors_count"), field(info, "indexing_threshold")); + } + t.write(); + } + + @Test @Order(9) + void redisFilterWithoutDeclaredSchema() { + assumeTrue(stores.containsKey(StoreFactory.REDIS)); + var t = new Transcript("09-redis-undeclared-metadata.txt"); + VectorStore vs = StoreFactory.redis(jedis, em, null, false); + ingest(vs, ds.documents().subList(0, 2_000)); + try { + Thread.sleep(1000); + List res = vs.similaritySearch(request(WARMUP, "category == 'c3'")); + t.line("metadata fields NOT declared in the builder; filter category == 'c3' -> %d results", res.size()); + t.line("unfiltered on the same index -> %d results", vs.similaritySearch(request(WARMUP, null)).size()); + } catch (Exception e) { + t.line("metadata fields NOT declared; filter category == 'c3' threw %s", e.getClass().getSimpleName()); + t.line(" %s", String.valueOf(e.getMessage()).lines().findFirst().orElse("")); + } + t.write(); + } + + + @Test @Order(10) + void elasticsearchIndexOptions() throws Exception { + assumeTrue(stores.containsKey(StoreFactory.ES)); + var t = new Transcript("10-elasticsearch-index-options.txt"); + t.line("same 30,000 documents, same queries; only the dense_vector index_options.type differs (the first row is Spring AI's own mapping)"); + t.line("%-12s %10s %9s %12s %14s %10s", "type", "recall@10", "p50 ms", "filt recall", "filt p50 ms", "store"); + Result d = measure(stores.get(StoreFactory.ES), null, false); + Result df = measure(stores.get(StoreFactory.ES), FILTER, true); + t.line("%-12s %10.3f %9.1f %12.3f %14.1f %10s", "bbq_hnsw*", d.recall, d.p50, df.recall, df.p50, esStoreSize(StoreFactory.NAME)); + for (String type : new String[] {"int8_hnsw", "hnsw"}) { + String index = "vs_es_" + type; + StoreFactory.http("DELETE", "http://127.0.0.1:9200/" + index, null); + StoreFactory.http("PUT", "http://127.0.0.1:9200/" + index, "{\"mappings\":{\"properties\":{\"embedding\":{\"type\":\"dense_vector\",\"dims\":384,\"index\":true,\"similarity\":\"cosine\",\"index_options\":{\"type\":\"" + type + "\"}}}}}"); + var o = new org.springframework.ai.vectorstore.elasticsearch.ElasticsearchVectorStoreOptions(); + o.setIndexName(index); + o.setDimensions(Dataset.DIMS); + o.setSimilarity(org.springframework.ai.vectorstore.elasticsearch.SimilarityFunction.cosine); + VectorStore vs = StoreFactory.init(org.springframework.ai.vectorstore.elasticsearch.ElasticsearchVectorStore.builder(es, em).options(o).initializeSchema(false).build()); + StoreFactory.http("GET", "http://127.0.0.1:9200/_cluster/health/" + index + "?wait_for_status=yellow&timeout=60s", null); + ingest(vs, ds.documents()); + StoreFactory.http("POST", "http://127.0.0.1:9200/" + index + "/_refresh", null); + Result r = measure(vs, null, false); + Result rf = measure(vs, FILTER, true); + t.line("%-12s %10.3f %9.1f %12.3f %14.1f %10s", type, r.recall, r.p50, rf.recall, rf.p50, esStoreSize(index)); + } + t.line("* default; its full mapping is in 04-defaults-that-matter.txt"); + t.write(); + } + + String esStoreSize(String index) throws Exception { + StoreFactory.http("POST", "http://127.0.0.1:9200/" + index + "/_forcemerge?max_num_segments=1", null); + return StoreFactory.http("GET", "http://127.0.0.1:9200/_cat/indices/" + index + "?h=store.size", null).trim(); + } + + /** The JSON object value of the first occurrence of "name":{...}, braces matched. */ + static String object(String json, String name) { + int i = json.indexOf("\"" + name + "\":{"); + if (i < 0) return "?"; + int start = json.indexOf('{', i), depth = 0; + for (int p = start; p < json.length(); p++) { + if (json.charAt(p) == '{') depth++; + if (json.charAt(p) == '}' && --depth == 0) return "\"" + name + "\":" + json.substring(start, p + 1); + } + return "?"; + } + + static String field(String json, String name) { + var m = java.util.regex.Pattern.compile("\"" + name + "\":\\s*([^,}]+)").matcher(json); + return m.find() ? m.group(1) : "?"; + } +} diff --git a/vector-stores/src/test/java/com/ankurm/vectorstores/support/Transcript.java b/vector-stores/src/test/java/com/ankurm/vectorstores/support/Transcript.java new file mode 100644 index 0000000..7877686 --- /dev/null +++ b/vector-stores/src/test/java/com/ankurm/vectorstores/support/Transcript.java @@ -0,0 +1,35 @@ +package com.ankurm.vectorstores.support; + +import java.io.IOException; +import java.nio.file.Files; +import java.nio.file.Path; +import java.util.ArrayList; +import java.util.List; + +/** Collects lines, echoes them, and writes output/NN-name.txt -- every console block in the article is quoted from one. */ +public final class Transcript { + + private final List lines = new ArrayList<>(); + private final String file; + + public Transcript(String file) { + this.file = file; + } + + public Transcript line(String fmt, Object... args) { + String s = args.length == 0 ? fmt : String.format(fmt, args); + lines.add(s); + System.out.println("[" + file + "] " + s); + return this; + } + + public void write() { + try { + Path p = Path.of("output", file); + Files.createDirectories(p.getParent()); + Files.write(p, lines); + } catch (IOException e) { + throw new IllegalStateException(e); + } + } +}