Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01G8ikz8xdWuTP5yun8DZ1hk
60 lines
4.2 KiB
Markdown
60 lines
4.2 KiB
Markdown
# semantic-cache
|
|
|
|
Companion code for [Semantic Caching for LLM Calls in Spring Boot with Redis Vector Search](https://ankurm.com/semantic-caching-for-llm-calls-in-spring-boot-with-redis-vector-search/), part of the [Spring AI series](../README.md) on ankurm.com.
|
|
|
|
A `CallAdvisor` that answers from Redis when a question is close enough to one answered before, and otherwise calls the model and remembers the answer. Embeddings come from a real model, all-MiniLM-L6-v2, run in-process; the cache is a Redis Stack vector index driven through Spring AI's `RedisVectorStore`. A labelled set of 30 support questions (four phrasings each) and 20 off-topic ones is used to measure hit rate against wrong answers at every threshold. Every figure in the article is quoted from a file in [`output/`](output).
|
|
|
|
The *chat model* in the measurements is a script that knows the right answer for each question and counts calls, so a wrong answer is detectable. Dollar figures use an **illustrative** price ($2.50 / $10.00 per million tokens) and tokens estimated as characters / 4. One test (`10-real-model-latency.txt`) puts a real local model, TinyLlama 1.1B on Ollama, behind the same advisor; it is skipped when no Ollama is listening.
|
|
|
|
## Versions
|
|
|
|
| Component | Version |
|
|
|---|---|
|
|
| Spring Boot | 4.1.1 |
|
|
| Spring AI | 2.0.1 (`spring-ai-redis-store`, `spring-ai-transformers`; `spring-ai-ollama` for the optional test) |
|
|
| Java | 25 (Temurin 25.0.4.1) |
|
|
| Redis Stack | 7.4.0-v8 tarball: Redis 7.4.7, RediSearch 2.10.20, RedisJSON 2.8.9 |
|
|
| Jedis | 7.4.1 |
|
|
| Embedding model | all-MiniLM-L6-v2, 384 dimensions, ONNX via ONNX Runtime 1.21.1 and DJL tokenizers 0.36.0 |
|
|
|
|
## Quickstart (no Docker)
|
|
|
|
```bash
|
|
scripts/services-up.sh # Redis Stack on :6393 (downloads the tarball if missing)
|
|
scripts/run-all.sh # runs the 10 tests and regenerates output/01 .. 10
|
|
```
|
|
|
|
The first run downloads the embedding model (about 87 MB) from Hugging Face into `$TMPDIR/minilm-cache`. RedisJSON must be loaded as well as RediSearch: `RedisVectorStore` stores JSON documents.
|
|
|
|
## What's here
|
|
|
|
| File | What it shows |
|
|
|---|---|
|
|
| [`SemanticCache.java`](src/main/java/com/ankurm/semanticcache/SemanticCache.java) | The Redis vector index behind one class: `lookup`, `put`, tenant filter, expiry |
|
|
| [`SemanticCacheAdvisor.java`](src/main/java/com/ankurm/semanticcache/SemanticCacheAdvisor.java) | The `CallAdvisor`: lookup, short-circuit on a hit, store on a miss |
|
|
| [`Embedder.java`](src/main/java/com/ankurm/semanticcache/Embedder.java) | MiniLM loaded from Hugging Face, plus a plain cosine for comparison |
|
|
| [`ScriptedModel.java`](src/main/java/com/ankurm/semanticcache/ScriptedModel.java) | A chat model that knows each question's right answer and counts calls and tokens |
|
|
| [`Dataset.java`](src/main/java/com/ankurm/semanticcache/Dataset.java), [`dataset.tsv`](src/main/resources/dataset.tsv) | 30 intents x 4 phrasings (including look-alike groups such as reset password / reset router / reset 2FA) and 20 off-topic questions |
|
|
| [`SemanticCacheTest.java`](src/test/java/com/ankurm/semanticcache/SemanticCacheTest.java) | Every measurement; each test writes its own transcript |
|
|
|
|
## Output files
|
|
|
|
| File | Written by |
|
|
|---|---|
|
|
| `01-similarity-scores.txt` | `similarityScoresOfParaphrasesAndLookAlikes` (also: Spring AI's score is `(1 + cosine) / 2`) |
|
|
| `02-advisor-in-a-chat-client.txt` | `theAdvisorServesAParaphraseWithoutCallingTheModel` |
|
|
| `03-threshold-sweep.txt` | `thresholdSweepHitRateAgainstWrongAnswers` |
|
|
| `04-replay-600-requests.txt` | `replayingSixHundredRequests` |
|
|
| `05-lookalikes.txt` | `lookAlikesThatShareWordsButNotMeaning` |
|
|
| `06-tenant-isolation.txt` | `tenantsDoNotShareAnswers` |
|
|
| `07-expiry.txt` | `anExpiredEntryStopsBeingServed` |
|
|
| `08-embedding-model-change.txt` | `anotherEmbeddingModelIsAnotherIndex` |
|
|
| `09-hit-latency.txt` | `whereTheTimeGoesOnAHit` |
|
|
| `10-real-model-latency.txt` | `againstARealLocalModel` (needs Ollama on :11555 with a model named `tl`) |
|
|
|
|
Counts and scores are deterministic; timings drift between runs (2-vCPU sandbox).
|
|
|
|
## Not covered
|
|
|
|
Streaming responses (the advisor is call-only), caching answers that depend on conversation history or tool results, cache warming, and an embedding model other than MiniLM: the thresholds here are for this model and this kind of short support question.
|