# semantic-cache Companion code for [Semantic Caching for LLM Calls in Spring Boot with Redis Vector Search](https://ankurm.com/semantic-caching-for-llm-calls-in-spring-boot-with-redis-vector-search/), part of the [Spring AI series](../README.md) on ankurm.com. A `CallAdvisor` that answers from Redis when a question is close enough to one answered before, and otherwise calls the model and remembers the answer. Embeddings come from a real model, all-MiniLM-L6-v2, run in-process; the cache is a Redis Stack vector index driven through Spring AI's `RedisVectorStore`. A labelled set of 30 support questions (four phrasings each) and 20 off-topic ones is used to measure hit rate against wrong answers at every threshold. Every figure in the article is quoted from a file in [`output/`](output). The *chat model* in the measurements is a script that knows the right answer for each question and counts calls, so a wrong answer is detectable. Dollar figures use an **illustrative** price ($2.50 / $10.00 per million tokens) and tokens estimated as characters / 4. One test (`10-real-model-latency.txt`) puts a real local model, TinyLlama 1.1B on Ollama, behind the same advisor; it is skipped when no Ollama is listening. ## Versions | Component | Version | |---|---| | Spring Boot | 4.1.1 | | Spring AI | 2.0.1 (`spring-ai-redis-store`, `spring-ai-transformers`; `spring-ai-ollama` for the optional test) | | Java | 25 (Temurin 25.0.4.1) | | Redis Stack | 7.4.0-v8 tarball: Redis 7.4.7, RediSearch 2.10.20, RedisJSON 2.8.9 | | Jedis | 7.4.1 | | Embedding model | all-MiniLM-L6-v2, 384 dimensions, ONNX via ONNX Runtime 1.21.1 and DJL tokenizers 0.36.0 | ## Quickstart (no Docker) ```bash scripts/services-up.sh # Redis Stack on :6393 (downloads the tarball if missing) scripts/run-all.sh # runs the 10 tests and regenerates output/01 .. 10 ``` The first run downloads the embedding model (about 87 MB) from Hugging Face into `$TMPDIR/minilm-cache`. RedisJSON must be loaded as well as RediSearch: `RedisVectorStore` stores JSON documents. ## What's here | File | What it shows | |---|---| | [`SemanticCache.java`](src/main/java/com/ankurm/semanticcache/SemanticCache.java) | The Redis vector index behind one class: `lookup`, `put`, tenant filter, expiry | | [`SemanticCacheAdvisor.java`](src/main/java/com/ankurm/semanticcache/SemanticCacheAdvisor.java) | The `CallAdvisor`: lookup, short-circuit on a hit, store on a miss | | [`Embedder.java`](src/main/java/com/ankurm/semanticcache/Embedder.java) | MiniLM loaded from Hugging Face, plus a plain cosine for comparison | | [`ScriptedModel.java`](src/main/java/com/ankurm/semanticcache/ScriptedModel.java) | A chat model that knows each question's right answer and counts calls and tokens | | [`Dataset.java`](src/main/java/com/ankurm/semanticcache/Dataset.java), [`dataset.tsv`](src/main/resources/dataset.tsv) | 30 intents x 4 phrasings (including look-alike groups such as reset password / reset router / reset 2FA) and 20 off-topic questions | | [`SemanticCacheTest.java`](src/test/java/com/ankurm/semanticcache/SemanticCacheTest.java) | Every measurement; each test writes its own transcript | ## Output files | File | Written by | |---|---| | `01-similarity-scores.txt` | `similarityScoresOfParaphrasesAndLookAlikes` (also: Spring AI's score is `(1 + cosine) / 2`) | | `02-advisor-in-a-chat-client.txt` | `theAdvisorServesAParaphraseWithoutCallingTheModel` | | `03-threshold-sweep.txt` | `thresholdSweepHitRateAgainstWrongAnswers` | | `04-replay-600-requests.txt` | `replayingSixHundredRequests` | | `05-lookalikes.txt` | `lookAlikesThatShareWordsButNotMeaning` | | `06-tenant-isolation.txt` | `tenantsDoNotShareAnswers` | | `07-expiry.txt` | `anExpiredEntryStopsBeingServed` | | `08-embedding-model-change.txt` | `anotherEmbeddingModelIsAnotherIndex` | | `09-hit-latency.txt` | `whereTheTimeGoesOnAHit` | | `10-real-model-latency.txt` | `againstARealLocalModel` (needs Ollama on :11555 with a model named `tl`) | Counts and scores are deterministic; timings drift between runs (2-vCPU sandbox). ## Not covered Streaming responses (the advisor is call-only), caching answers that depend on conversation history or tool results, cache warming, and an embedding model other than MiniLM: the thresholds here are for this model and this kind of short support question.