Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01G8ikz8xdWuTP5yun8DZ1hk
semantic-cache
Companion code for Semantic Caching for LLM Calls in Spring Boot with Redis Vector Search, part of the Spring AI series on ankurm.com.
A CallAdvisor that answers from Redis when a question is close enough to one answered before, and otherwise calls the model and remembers the answer. Embeddings come from a real model, all-MiniLM-L6-v2, run in-process; the cache is a Redis Stack vector index driven through Spring AI's RedisVectorStore. A labelled set of 30 support questions (four phrasings each) and 20 off-topic ones is used to measure hit rate against wrong answers at every threshold. Every figure in the article is quoted from a file in output/.
The chat model in the measurements is a script that knows the right answer for each question and counts calls, so a wrong answer is detectable. Dollar figures use an illustrative price ($2.50 / $10.00 per million tokens) and tokens estimated as characters / 4. One test (10-real-model-latency.txt) puts a real local model, TinyLlama 1.1B on Ollama, behind the same advisor; it is skipped when no Ollama is listening.
Versions
| Component | Version |
|---|---|
| Spring Boot | 4.1.1 |
| Spring AI | 2.0.1 (spring-ai-redis-store, spring-ai-transformers; spring-ai-ollama for the optional test) |
| Java | 25 (Temurin 25.0.4.1) |
| Redis Stack | 7.4.0-v8 tarball: Redis 7.4.7, RediSearch 2.10.20, RedisJSON 2.8.9 |
| Jedis | 7.4.1 |
| Embedding model | all-MiniLM-L6-v2, 384 dimensions, ONNX via ONNX Runtime 1.21.1 and DJL tokenizers 0.36.0 |
Quickstart (no Docker)
scripts/services-up.sh # Redis Stack on :6393 (downloads the tarball if missing)
scripts/run-all.sh # runs the 10 tests and regenerates output/01 .. 10
The first run downloads the embedding model (about 87 MB) from Hugging Face into $TMPDIR/minilm-cache. RedisJSON must be loaded as well as RediSearch: RedisVectorStore stores JSON documents.
What's here
| File | What it shows |
|---|---|
SemanticCache.java |
The Redis vector index behind one class: lookup, put, tenant filter, expiry |
SemanticCacheAdvisor.java |
The CallAdvisor: lookup, short-circuit on a hit, store on a miss |
Embedder.java |
MiniLM loaded from Hugging Face, plus a plain cosine for comparison |
ScriptedModel.java |
A chat model that knows each question's right answer and counts calls and tokens |
Dataset.java, dataset.tsv |
30 intents x 4 phrasings (including look-alike groups such as reset password / reset router / reset 2FA) and 20 off-topic questions |
SemanticCacheTest.java |
Every measurement; each test writes its own transcript |
Output files
| File | Written by |
|---|---|
01-similarity-scores.txt |
similarityScoresOfParaphrasesAndLookAlikes (also: Spring AI's score is (1 + cosine) / 2) |
02-advisor-in-a-chat-client.txt |
theAdvisorServesAParaphraseWithoutCallingTheModel |
03-threshold-sweep.txt |
thresholdSweepHitRateAgainstWrongAnswers |
04-replay-600-requests.txt |
replayingSixHundredRequests |
05-lookalikes.txt |
lookAlikesThatShareWordsButNotMeaning |
06-tenant-isolation.txt |
tenantsDoNotShareAnswers |
07-expiry.txt |
anExpiredEntryStopsBeingServed |
08-embedding-model-change.txt |
anotherEmbeddingModelIsAnotherIndex |
09-hit-latency.txt |
whereTheTimeGoesOnAHit |
10-real-model-latency.txt |
againstARealLocalModel (needs Ollama on :11555 with a model named tl) |
Counts and scores are deterministic; timings drift between runs (2-vCPU sandbox).
Not covered
Streaming responses (the advisor is call-only), caching answers that depend on conversation history or tool results, cache warming, and an embedding model other than MiniLM: the thresholds here are for this model and this kind of short support question.