Files

4.2 KiB

semantic-cache

Companion code for Semantic Caching for LLM Calls in Spring Boot with Redis Vector Search, part of the Spring AI series on ankurm.com.

A CallAdvisor that answers from Redis when a question is close enough to one answered before, and otherwise calls the model and remembers the answer. Embeddings come from a real model, all-MiniLM-L6-v2, run in-process; the cache is a Redis Stack vector index driven through Spring AI's RedisVectorStore. A labelled set of 30 support questions (four phrasings each) and 20 off-topic ones is used to measure hit rate against wrong answers at every threshold. Every figure in the article is quoted from a file in output/.

The chat model in the measurements is a script that knows the right answer for each question and counts calls, so a wrong answer is detectable. Dollar figures use an illustrative price ($2.50 / $10.00 per million tokens) and tokens estimated as characters / 4. One test (10-real-model-latency.txt) puts a real local model, TinyLlama 1.1B on Ollama, behind the same advisor; it is skipped when no Ollama is listening.

Versions

Component Version
Spring Boot 4.1.1
Spring AI 2.0.1 (spring-ai-redis-store, spring-ai-transformers; spring-ai-ollama for the optional test)
Java 25 (Temurin 25.0.4.1)
Redis Stack 7.4.0-v8 tarball: Redis 7.4.7, RediSearch 2.10.20, RedisJSON 2.8.9
Jedis 7.4.1
Embedding model all-MiniLM-L6-v2, 384 dimensions, ONNX via ONNX Runtime 1.21.1 and DJL tokenizers 0.36.0

Quickstart (no Docker)

scripts/services-up.sh     # Redis Stack on :6393 (downloads the tarball if missing)
scripts/run-all.sh         # runs the 10 tests and regenerates output/01 .. 10

The first run downloads the embedding model (about 87 MB) from Hugging Face into $TMPDIR/minilm-cache. RedisJSON must be loaded as well as RediSearch: RedisVectorStore stores JSON documents.

What's here

File What it shows
SemanticCache.java The Redis vector index behind one class: lookup, put, tenant filter, expiry
SemanticCacheAdvisor.java The CallAdvisor: lookup, short-circuit on a hit, store on a miss
Embedder.java MiniLM loaded from Hugging Face, plus a plain cosine for comparison
ScriptedModel.java A chat model that knows each question's right answer and counts calls and tokens
Dataset.java, dataset.tsv 30 intents x 4 phrasings (including look-alike groups such as reset password / reset router / reset 2FA) and 20 off-topic questions
SemanticCacheTest.java Every measurement; each test writes its own transcript

Output files

File Written by
01-similarity-scores.txt similarityScoresOfParaphrasesAndLookAlikes (also: Spring AI's score is (1 + cosine) / 2)
02-advisor-in-a-chat-client.txt theAdvisorServesAParaphraseWithoutCallingTheModel
03-threshold-sweep.txt thresholdSweepHitRateAgainstWrongAnswers
04-replay-600-requests.txt replayingSixHundredRequests
05-lookalikes.txt lookAlikesThatShareWordsButNotMeaning
06-tenant-isolation.txt tenantsDoNotShareAnswers
07-expiry.txt anExpiredEntryStopsBeingServed
08-embedding-model-change.txt anotherEmbeddingModelIsAnotherIndex
09-hit-latency.txt whereTheTimeGoesOnAHit
10-real-model-latency.txt againstARealLocalModel (needs Ollama on :11555 with a model named tl)

Counts and scores are deterministic; timings drift between runs (2-vCPU sandbox).

Not covered

Streaming responses (the advisor is call-only), caching answers that depend on conversation history or tool results, cache warming, and an embedding model other than MiniLM: the thresholds here are for this model and this kind of short support question.