Skip to main content

Semantic Caching for LLM Calls in Spring Boot with Redis Vector Search

Build a semantic cache for LLM calls in Spring Boot with Spring AI and Redis vector search, then measure threshold, hit rate, wrong answers and savings on 600 replayed requests.

A call to a large language model is slow and it costs money, and a surprising share of the questions it gets are the same question in different words. “How do I reset my password?”, “I forgot my password, how do I get back in?” and “steps to change a lost password” all deserve the same answer. A normal cache cannot see that, because it matches the text exactly. A semantic cache matches the meaning: it stores each answer under a numeric fingerprint of the question and, for a new question, looks for a stored fingerprint that is close enough. This article builds one in Spring Boot 4 with Spring AI 2.0, using Redis as the store, and then does the part most tutorials skip: it measures it. You will see what “close enough” looks like as numbers, how the threshold trades hits against wrong answers, what the cache saves on 600 replayed requests, and several ways it goes wrong. Depth is in expandable sections, so you can read straight through or open only what you need. If you want the Redis basics first, Redis with Spring Boot 4.1 covers them.
Versions, and honest limits. Spring Boot 4.1.1, Spring AI 2.0.1 (spring-ai-redis-store, spring-ai-transformers), Redis Stack 7.4.0-v8 (Redis 7.4.7, RediSearch 2.10.20, RedisJSON 2.8.9), Jedis 7.4.1, the embedding model all-MiniLM-L6-v2 (384 dimensions) run in-process, and Java 25. All the code is in the semantic-cache module of asmhatre/spring-ai, and every console block below is quoted from a file under semantic-cache/output/, written by a test. Three things to know before trusting any number. The 30 questions and their rewordings are ones I wrote, so real users will vary more than my paraphrases do. The chat model in the measurements is a script that knows the right answer to each question and counts calls, so that a wrong answer can be detected; dollar figures use an illustrative price ($2.50 and $10.00 per million tokens) and tokens counted as characters divided by four. One test, near the end, does use a real model. And all timings come from a cloud sandbox with 2 virtual CPUs, not a laptop or a production server.

Why a normal cache misses most repeated questions

Put the request text in a map and you have a cache. It works when the same text comes back, and that is rarer than it sounds, because people phrase things their own way. The semantic version adds two steps in front of the lookup: turn the question into a list of numbers that captures its meaning (an embedding), then ask a vector index for the stored question whose numbers are nearest. If the nearest one is close enough, return its stored answer and skip the model. If not, call the model and store the new answer for next time.
Questionany wordingEmbedtext to 384 numbersRedis searchnearest stored questionClose enough?score above thresholdAnswercached, or ask the model
The diagram has one decision in it, and it is the whole subject of this article. Set the bar too high and the cache only catches near-identical questions. Set it too low and it answers a different question with a stored answer, and nobody sees an error when that happens. The measurements below are about where to put that bar.

Turning a question into numbers you can compare

An embedding model reads a piece of text and returns a fixed-length list of numbers, so that texts with similar meaning get lists that point in similar directions. The model here is all-MiniLM-L6-v2, a small public one that returns 384 numbers per text. Spring AI’s spring-ai-transformers module runs it inside your JVM, so there is no embedding API to call or pay for. By default it downloads the model from a GitHub media URL; I pointed it at Hugging Face and gave it a cache directory, so the 87 MB download happens once. From Embedder.java:
TransformersEmbeddingModel m = new TransformersEmbeddingModel();
m.setTokenizerResource(HF + "tokenizer.json");
m.setModelResource(HF + "onnx/model.onnx");
m.setResourceCacheDirectory(System.getProperty("java.io.tmpdir") + "/minilm-cache");
m.setTokenizerOptions(java.util.Map.of("padding", "true", "truncation", "true", "maxLength", "256"));
m.afterPropertiesSet();
To compare two of these lists, the usual measure is cosine similarity: 1.0 means pointing the same way, 0 means unrelated. The first test embeds the 30 cached questions and 90 rewordings of them and reports how the scores spread, in 01-similarity-scores.txt:
Plain cosine similarity (NOT the Spring AI score, see the end) of each later phrasing with the cached phrasing of the SAME question (90 pairs):
  min 0.337   p10 0.537   median 0.729   p90 0.869   max 0.916
  lowest: 0.337   Where is my package right now?  ->  How can I track my order?

Best score each of those phrasings gets against a cached question of a DIFFERENT intent (look-alikes like reset password / reset router):
  min 0.209   p10 0.311   median 0.446   p90 0.563   max 0.633
  highest: 0.633   Can you help me reset the password for my login?  ->  How do I reset two-factor authentication on my account?

Best score of 20 off-topic questions against any cached question:
  min 0.096   p10 0.133   median 0.180   p90 0.225   max 0.289
Read the three groups as three bands. A reworded question lands anywhere from 0.34 to 0.92 against its own cached original, with a median of 0.73. A question about a different topic in the same shop (the best match of each reworded question against everything else that is cached) reaches at most 0.63. Off-topic questions top out at 0.29. So the bands are separate, but not by much: the lowest reworded question, “Where is my package right now?” against “How can I track my order?”, scores 0.34, which is below the best look-alike. No threshold will catch every rewording and reject every different question, and the next sections put numbers on that trade.
Going deeper: what I put in the dataset
The questions are in dataset.tsv: 30 intents for an imaginary electronics shop, each with one first phrasing and three rewordings, plus 20 off-topic questions. Several intents are deliberately close neighbours so that a too-generous threshold gets caught out: reset a password, a router or two-factor authentication; cancel an order, a subscription or a return; track an order or a return; change an email, an address or a phone number. The label on each row is its intent, so a served answer can be marked right (same intent as the question) or wrong (another intent). Dataset.java loads it. I wrote all of it, and an author’s rewordings are tidier than real users’ typos, abbreviations and half-sentences. Treat the hit rates as what this model does on clean rewordings, an upper bound for messy traffic.

The score Spring AI gives you is not the cosine

Here is a detail that changes how you set the threshold. The same test asks Redis, through Spring AI, for the nearest stored question to one reworded question, and also computes the cosine in Java. They disagree. From 01-similarity-scores.txt:
Does Redis report the same number? "I forgot my password, how can I get back into my account?"
  cosine computed in Java        : 0.838726
  Document.getScore() from Redis : 0.919363   (matched stored question: "How do I reset my account password?")
  (1 + cosine) / 2               : 0.919363   <- the score is this, not the cosine
  same question as the cached one -> score 1.000000

Spring AI's similarityThreshold is applied to that score. Threshold -> the cosine it really means (cosine = 2 x score - 1):
  similarityThreshold 0.70  =  cosine 0.40
  similarityThreshold 0.80  =  cosine 0.60
  similarityThreshold 0.85  =  cosine 0.70
  similarityThreshold 0.90  =  cosine 0.80
  similarityThreshold 0.95  =  cosine 0.90
The score on a returned Document is (1 + cosine) / 2, not the cosine: 0.8387 in Java became 0.9194 from Redis. Spring AI uses that score for SearchRequest.similarityThreshold too, so a threshold of 0.80 does not mean “80% similar”. It means a cosine of 0.60, which in the bands above is already a generous bar. Every threshold in the rest of this article is on the Spring AI score, the number you would type into SearchRequest, and the sweep table has a cosine column beside it so you can translate.
A trap with no error message. If you copy a threshold from a blog post, or from a library whose scores are plain cosine, it will mean something quite different here. This is the first thing to check, and the quickest check is the one above: embed two sentences, compute the cosine yourself, and compare it with the score the store returns.

Redis as the cache

Redis Stack adds a search module to Redis that can hold vectors and find the nearest ones. Spring AI wraps it as a RedisVectorStore, so storing a question is adding a Document and finding the nearest one is a similarity search. The class that holds it, SemanticCache.java, builds the store by hand so the index settings are visible:
        this.store = RedisVectorStore.builder(jedis, embeddings)
                .indexName(INDEX)
                .prefix(PREFIX)
                .vectorAlgorithm(RedisVectorStore.Algorithm.HNSW)
                .distanceMetric(RedisVectorStore.DistanceMetric.COSINE)
                // Only declared metadata fields are filterable, and only declared ones come back on a search result:
                // without text("answer") a hit has the stored question but no answer.
                .metadataFields(RedisVectorStore.MetadataField.tag("tenant"), RedisVectorStore.MetadataField.text("answer"))
                .initializeSchema(true)
                .build();
        this.store.afterPropertiesSet();
Three details in that block are worth knowing. The store keeps each entry as a JSON document, so it needs RedisJSON as well as the search module; with only the search module loaded, creating the index fails with “Invalid rule type: JSON”, which is how I found out. The distance is cosine, and the vector index is HNSW. And only metadata fields you declare can be filtered on, and only declared fields come back on a search result: my first version stored the answer as undeclared metadata and every hit returned a null answer. A lookup is then one search: the question, the nearest one result, the threshold, and a filter on the tenant. The filter is explained in a later section. From SemanticCache.java:
    public Optional<Hit> lookup(String tenant, String question, double minScore) {
        List<Document> found = store.similaritySearch(SearchRequest.builder()
                .query(question)
                .topK(1)
                .similarityThreshold(minScore)
                .filterExpression("tenant == '" + tenant + "'")
                .build());
        return found.stream().findFirst().map(d -> new Hit(
                (String) d.getMetadata().get("answer"), d.getText(), d.getScore()));
    }
Going deeper: what declaring the answer as a field costs
MetadataField.text("answer") makes Redis build a full-text index over every answer, which a cache never searches. It is the simplest way to get the answer back, and fine at this size, but it is wasted index memory. The alternative is to put only the question and a key in the vector document and keep the answer in a separate Redis string under the same id, with the same expiry. I did not build or measure that variant. The store also needs Redis Stack, or Redis 8, which includes the modules. The service script services-up.sh downloads the Redis Stack tarball and starts it with both modules, with no Docker.

The advisor that sits in front of the model

In Spring AI, an advisor is a step that sees every ChatClient request before and after the model. A semantic cache is a natural advisor: look up first, and if there is a hit, return without calling the rest of the chain. From SemanticCacheAdvisor.java:
    public ChatClientResponse adviseCall(ChatClientRequest request, CallAdvisorChain chain) {
        String tenant = tenantOf.apply(request);
        String question = request.prompt().getUserMessage().getText();

        Optional<SemanticCache.Hit> hit = cache.lookup(tenant, question);
        if (hit.isPresent()) {
            hits.incrementAndGet();
            ChatResponse cached = ChatResponse.builder()
                    .generations(java.util.List.of(new Generation(new AssistantMessage(hit.get().answer()))))
                    .build();
            return new ChatClientResponse(cached, Map.of(HIT, true, SCORE, hit.get().score(), MATCHED, hit.get().storedQuestion()));
        }

        misses.incrementAndGet();
        ChatClientResponse response = chain.nextCall(request);
        String answer = response.chatResponse().getResult().getOutput().getText();
        if (answer != null && !answer.isBlank()) {
            cache.put(tenant, question, answer);
        }
        return response;
    }
On a hit the advisor builds a ChatResponse itself and returns it, so everything later in the chain, including the call to the model, never runs. On a miss it lets the call through and stores the answer. It puts a flag and the score in the response context so a caller can tell a cached answer from a fresh one. Five questions through a client with the advisor at threshold 0.80, from 02-advisor-in-a-chat-client.txt:
# SemanticCacheAdvisor in front of a ChatClient (threshold 0.80)

ask : How do I reset my account password?
  -> model call, model calls so far 1, answer starts "[reset-password]"
ask : How do I reset my account password?
  -> CACHE HIT, score 1.000, model calls so far 1, answer starts "[reset-password]"
ask : I forgot my password, how can I get back into my account?
  -> CACHE HIT, score 0.919, model calls so far 1, answer starts "[reset-password]"
ask : How do I reset my router to factory settings?
  -> model call, model calls so far 2, answer starts "[reset-router]"
ask : What is the capital of Australia?
  -> model call, model calls so far 3, answer starts "[ood-0]"

advisor counters: 2 hits, 3 misses; model received 3 calls for 5 questions; entries in Redis: 3
Asking the identical question again is a hit with score 1.000. The reworded password question is a hit at 0.919, and the model was not called. The router question and the unrelated one went through to the model. Five questions, three model calls. One placement note: register the advisor with an early order, ahead of anything that adds cost or side effects, such as tool calling or chat memory, because everything after it is skipped on a hit.
Going deeper: where it sits relative to other advisors, and what it skips
An advisor chain runs in order of getOrder(), lowest first on the way in. The cache here takes whatever order you give it, and a hit returns immediately, so advisors with a higher order never see that request: a logging advisor placed after it will not log cached answers, and a token-budget advisor after it will not count them. That may be what you want, but decide it on purpose. The advisors article, Writing Custom Advisors in Spring AI 2.0, covers ordering in detail. This advisor implements only CallAdvisor, so a streaming request bypasses it entirely. A cached answer could be streamed back as one chunk, but I did not write that. A request whose answer depends on earlier turns of a conversation, or on a tool result, is also a poor fit: the cache keys on the last user message alone. SemanticCacheAdvisor.java has the full class.

Choosing the threshold

Now the measurement the diagram promised. The test caches the 30 first phrasings, then asks each of the 90 rewordings and 20 off-topic questions, takes the nearest stored question, and counts, for each threshold, how many it would serve and whether the answer belongs to the question asked. From 03-threshold-sweep.txt:
threshold = the value given to SearchRequest.similarityThreshold (Spring AI score); cosine = 2 x threshold - 1.
A served answer is RIGHT if it belongs to the intent of the question asked, WRONG if it belongs to another intent.
An off-topic question has no right answer in the cache, so any hit on it is WRONG.

threshold | cosine | hit rate |     right | wrong (reworded) | wrong (off-topic) | wrong / all hits | cached pairs of different intents that answer for each other
     0.50 |   0.00 |     100% |    86/90  |        4/90     |        20/20      | 21.8% | 396 of 435
     0.60 |   0.20 |     100% |    86/90  |        4/90     |         6/20      | 10.4% | 151 of 435
     0.70 |   0.40 |      99% |    86/90  |        3/90     |         0/20      | 3.4% | 27 of 435
     0.75 |   0.50 |      94% |    83/90  |        2/90     |         0/20      | 2.4% | 13 of 435
     0.80 |   0.60 |      86% |    77/90  |        0/90     |         0/20      | 0.0% | 2 of 435
     0.85 |   0.70 |      58% |    52/90  |        0/90     |         0/20      | 0.0% | 1 of 435
     0.90 |   0.80 |      26% |    23/90  |        0/90     |         0/20      | 0.0% | 0 of 435
     0.95 |   0.90 |       3% |     3/90  |        0/90     |         0/20      | 0.0% | 0 of 435
Read the rows from the bottom. At 0.95 the cache is safe but catches 3% of rewordings, almost useless. At 0.90 it catches 26% with no wrong answers; at 0.85, 58%; at 0.80, 86%, still with no wrong answers on these 110 questions. Below that the wrong answers start: at 0.70, 3 of 90 rewordings got another question’s answer and none of the off-topic ones, and at 0.50 every off-topic question was answered from the cache, with an answer from some unrelated topic. The best operating point on this table looks like 0.80. The last column says why that is too optimistic. It counts pairs among the 30 cached questions themselves, 435 pairs, that are close enough to answer for each other. Two pairs qualify at 0.80 even though this table found no wrong answer, because the table only tested rewordings against the cache, not first-time questions that happen to sit next to a cached one. The closest pairs, and then a real model, show what that looks like.
The closest pairs of DIFFERENT cached questions (Spring AI score; 435 pairs in all). Above the threshold, one answers for the other:
  0.852  "How do I reset my account password?"  /  "How do I reset two-factor authentication on my account?"
  0.824  "How do I cancel an order I just placed?"  /  "How do I cancel a return I already requested?"
  0.796  "How do I cancel an order I just placed?"  /  "How do I cancel my monthly subscription?"
  0.791  "How can I track my order?"  /  "How can I track the return I sent back?"
  0.789  "How can I track my order?"  /  "Can I pick up my order at a store?"
  0.779  "How long is the warranty on your laptops?"  /  "How do I make a warranty claim?"
  0.775  "How long does a refund take to arrive?"  /  "How can I track the return I sent back?"
  0.767  "How do I update the firmware on my headphones?"  /  "How do I pair my headphones over Bluetooth?"
“Reset my password” and “reset my two-factor authentication” score 0.852, above the 0.80 threshold and above the 0.85 one. Someone who loses their phone and asks about two-factor would be told how to reset their password. That is not a rare edge case; in this dataset it is the closest pair of all, and I wrote the dataset to be a realistic shop. The test that produces this table is SemanticCacheTest.java.
Going deeper: the rewordings that missed at 0.80
Thirteen of the 90 rewordings scored below 0.80 and went to the model, which is the cost of the bar. From 03-threshold-sweep.txt:
Rewordings the cache MISSED at threshold 0.80 (best stored question and its score):
  0.787  "Where can I end my membership plan?"  (nearest: "How do I cancel my monthly subscription?")
  0.746  "Can you tell me what qualifies for getting my money back?"  (nearest: "Which items can be refunded?")
  0.696  "Where is my package right now?"  (nearest: "Can I pick up my order at a store?")
  0.759  "I want to see the delivery status of my purchase"  (nearest: "How can I track my order?")
  0.752  "Has my returned parcel arrived at your warehouse yet?"  (nearest: "How long will delivery take?")
  0.790  "My device broke, how do I get it repaired under warranty?"  (nearest: "How do I make a warranty claim?")
  0.797  "How many days until my parcel gets here?"  (nearest: "How long will delivery take?")
  0.739  "Can I pay with a credit card or PayPal?"  (nearest: "Which payment methods do you accept?")
  0.798  "How can I pay for my order?"  (nearest: "How can I track my order?")
  0.769  "I need a receipt for my purchase, where do I find it?"  (nearest: "Where can I download an invoice for my order?")
  0.743  "Is in-store collection available?"  (nearest: "Can I pick up my order at a store?")
  0.753  "I would rather collect my purchase myself, is that an option?"  (nearest: "Can I pick up my order at a store?")
  0.703  "Do you offer click and collect?"  (nearest: "Do you sell gift cards?")
Most are plausible misses (“Where can I end my membership plan?” scored 0.787 against the subscription question), and two are worse than a miss: “How can I pay for my order?” at 0.798 and “Do you offer click and collect?” at 0.703 were nearest to a different topic. Raising recall without admitting wrong answers is a job for a better embedding model or a reranking step, not for a lower threshold.

What it saves on 600 requests

A cache is judged on traffic, not on a table of pairs. The next test builds a workload of 600 requests with a fixed random seed: the 30 intents asked with a skewed popularity (the first about thirty times as often as the last), each time in a random one of its four phrasings, with 15% off-topic questions mixed in. It sends the same 600 through a plain client with no cache, and through the advisor at five settings. From 04-replay-600-requests.txt:
Workload: 600 requests (seed 42), 30 intents asked with Zipf popularity in random phrasings, 15% off-topic.
  121 distinct texts, so a plain exact-match cache could answer at most 479 of 600 (80%).
Cost model (ILLUSTRATIVE): $2.50 per million input tokens, $10.00 per million output tokens, tokens = characters / 4.

setup                  | model calls | hit rate |  wrong | wrong of hits |   model $ | saved
no cache               |        600 |        - |      - |             - |    0.4564 | -
cache, threshold 0.70  |         42 |    93.0% |    191 |         34.2% |    0.0316 | 93%
cache, threshold 0.80  |         62 |    89.7% |      7 |          1.3% |    0.0468 | 90%
cache, threshold 0.90  |        109 |    81.8% |      0 |          0.0% |    0.0827 | 82%
cache, threshold 0.95  |        119 |    80.2% |      0 |          0.0% |    0.0904 | 80%
cache, exact match only |        121 |    79.8% |      0 |          0.0% |    0.0919 | 80%

Average time to serve a cache hit (embed the question, search Redis, build the response): 4 ms at 0.70, 4 ms at 0.80, 3 ms at 0.90, 3 ms at 0.95
Saved dollars exclude the cost of embedding (local model here, CPU only) and of running Redis.
Look first at the row nobody builds: “exact match only”. A cache that compares plain text already answers 79.8% of this workload, because 600 requests produced only 121 distinct texts, and it is never wrong. The semantic cache at 0.90 adds two points on top (81.8%) with no wrong answers. At 0.80 it adds ten (89.7%) and serves 7 wrong answers among its 538 cache hits. At 0.70 it answers 93.0% of requests and 191 of those answers, 34%, are for a different question. That is the honest summary of what semantic caching buys on this workload: the big saving comes from repeats, which an ordinary cache gets too, and the semantic part is a smaller gain that is bought with wrong answers once the threshold drops. The dollar column is illustrative (tokens estimated as characters over four, a made-up price). It shows proportions, not your bill, and it leaves out the cost of embedding and of running Redis.
Going deeper: why these numbers depend on the workload I built
The 80% ceiling for exact matching is a result of my workload: 30 intents, four phrasings each, popularity skewed so that a few intents dominate. Real traffic with a longer tail of one-off questions has far fewer repeats, and the semantic part would then matter more or less depending on how many of the long-tail questions are rewordings of cached ones. The only way to know is to log your own traffic. The 15% of off-topic questions are never answered correctly from the cache at any sensible threshold; they all go to the model, which is correct behaviour. Seed 42 is fixed so the table regenerates identically; the counts do not drift between runs, but the millisecond figures in the same file do.

What a hit costs, against a real model

Two last measurements, about time. First, where the milliseconds go on a lookup. From 09-hit-latency.txt:
embed the question (MiniLM in-process) : median 2.4 ms, p95 3.6 ms
Redis KNN search with the tenant filter : median 0.3 ms, p95 2.5 ms   (lookup time minus one embedding)
Embedding the question takes about 2.4 ms, searching Redis about 0.3 ms. The embedding dominates, the search barely registers at 30 entries, and the whole hit costs a few milliseconds. Second, the same advisor in front of a real local model, TinyLlama 1.1B on Ollama, limited to 60 tokens. This is a very small model on two CPU cores, so its latency says nothing about a hosted frontier model, only that a model call is a different order of magnitude from a cache hit. From 10-real-model-latency.txt:
# A real local model (TinyLlama 1.1B, Q4_0, 60 tokens) behind the advisor, 2 vCPUs

first ask     2100 ms  model    "How do I reset my account password?"
first ask     1953 ms  model    "How do I reset my router to factory settings?"
first ask        5 ms  CACHE    "How do I reset two-factor authentication on my account?"  <- served the answer to "How do I reset my account password?" (score 0.852), a DIFFERENT question
first ask     1902 ms  model    "How do I cancel an order I just placed?"
first ask     1920 ms  model    "How do I cancel my monthly subscription?"
first ask       10 ms  CACHE    "How do I cancel a return I already requested?"  <- served the answer to "How do I cancel an order I just placed?" (score 0.824), a DIFFERENT question
reworded        17 ms  CACHE     "I forgot my password, how can I get back into my account?"  <- How do I reset my account password?
reworded         9 ms  CACHE     "What is the way to restore my router to its factory defaults?"  <- How do I reset my router to factory settings?
reworded      1941 ms  model     "I lost my phone, how can I reset my 2FA?"
reworded         9 ms  CACHE     "I made an order by mistake, can I cancel it?"  <- How do I cancel an order I just placed?
reworded         6 ms  CACHE     "I want to stop my recurring subscription payments"  <- How do I cancel my monthly subscription?
reworded      2013 ms  model     "I changed my mind about sending the item back, can I withdraw the return?"

median model call 1920 ms; median cache hit 9 ms (4 of 6 rewordings hit); first-time questions wrongly answered from the cache: 2 of 6
A model call took around 1.9 to 2.1 seconds; a hit took 5 to 17 milliseconds. Four of the six rewordings were hits. But look at the two lines marked CACHE under “first ask”: these were questions asked for the first time, and the cache answered both with the answer to a different question. “How do I reset two-factor authentication on my account?” got the password answer (score 0.852), and “How do I cancel a return I already requested?” got the answer for cancelling an order (0.824). The threshold was 0.80. Two of the six first-time questions were answered wrongly, instantly, with no error anywhere, and the real model never had the chance.

More ways it goes wrong

The look-alike pair above is a failure of the threshold. Four others are failures of what the embedding can see, who is asking, how long an answer stays true, and which model made the vectors. Meaning that differs in one word. Embeddings measure topical similarity, and negation, numbers and entity names change meaning without changing topic much. From 05-lookalikes.txt:
# Pairs that read alike and mean different things (cosine similarity)

cached question                               | new question                                  | cosine | Spring AI score = (1 + cosine) / 2
How do I cancel my order?                     | How do I keep my order and not cancel it?     |  0.938 | 0.969 <- served at threshold 0.90 and 0.80
Is shipping free for orders over $50?         | Is shipping free for orders over $500?        |  0.848 | 0.924 <- served at threshold 0.90 and 0.80
Do you ship to Canada?                        | Do you ship to Cuba?                          |  0.645 | 0.822 <- served at threshold 0.80
What is the warranty on the laptop?           | What is the warranty on the monitor?          |  0.799 | 0.899 <- served at threshold 0.80
Where is order 1001?                          | Where is order 2002?                          |  0.593 | 0.796 
Can I get a refund within 30 days?            | Can I get a refund after 30 days?             |  0.984 | 0.992 <- served at threshold 0.90 and 0.80
I want to delete my account                   | I do not want to delete my account            |  0.931 | 0.966 <- served at threshold 0.90 and 0.80
“Can I get a refund within 30 days?” and “Can I get a refund after 30 days?” score 0.992. “I want to delete my account” against “I do not want to delete my account” scores 0.966. Both would be served at a threshold of 0.90, and the second answer tells the user how to delete the account they said they want to keep. A model-based check on the match, or a cache restricted to questions where this cannot matter, is the only real defence; a higher threshold barely helps with 0.99.
Going deeper: tenants, expiry and changing the embedding model
Whose answer is it? If several customers share one cache, one customer’s answer can be served to another, including personalised ones. The tenant is a tag on each document and an expression on each search, so a search only sees its own tenant’s entries. 06-tenant-isolation.txt:
# Same question, two tenants

shop-a asks "How do I reset my account password?"  -> model call (calls so far: 1)
shop-b asks the identical question -> model call, no hit (model calls so far: 2)
shop-a asks again               -> cache hit
entries in the index: 2 (one per tenant)
The filter is a tag on the document and an expression on the search: tenant == 'shop-b'.
nearest(shop-a, q) without asking as shop-b still finds: "How do I reset my account password?"
The second tenant asking the identical question got no hit and caused a second model call, as it should. The two answers sit in the index side by side. The tenant here comes from a request-context parameter and defaults to default when not set, so a forgotten parameter silently shares one cache; the safer design makes the tenant mandatory. How long is an answer true? A cached answer about prices or opening hours goes stale. Redis can expire a key, and the index drops the document with it. 07-expiry.txt:
# Expiry with EXPIRE on the document's key

stored one entry, index holds 1 document(s); lookup -> hit
EXPIRE sc:<id> 1, wait 1.5 s; index now holds 0 document(s); lookup -> miss
What if you change the embedding model? Vectors from one model mean nothing to another, and different models have different lengths. The test builds the index with MiniLM (384 numbers) and then writes to the same index with an 8-number model. 08-embedding-model-change.txt:
# An index built for 384 dimensions, written by a model with 8

index created by MiniLM (384 dimensions); a new application version configures an 8-dimension model against the same index name.
add + search -> exception JedisDataException: Error parsing vector similarity query: query vector blob size (32) does not match index's expected size (1536).
With a different length, Redis refuses loudly, which is the good case. A new model with the same length, a newer version of the same model for example, would write vectors that fit the index and quietly mean something different. I did not test that case; the safe practice is to put the model name in the index name and let the old index expire.

Should you even do this?

A fair answer. Start with an exact-match cache on the normalised question (trim, lower-case, collapse whitespace). On this workload it caught 79.8% of requests with no wrong answers, and the semantic layer on top added two points at a safe threshold. A semantic cache is worth its extra moving parts (an embedding model, a vector index, a threshold to maintain, a class of silent errors) when your traffic is dominated by many phrasings of a small set of questions, the answers are public and identical for everyone, the cost of a wrong answer is low, and the answers do not go stale quickly: an FAQ assistant, say. It is a poor fit when answers depend on the user, the conversation or live data, or when a wrong answer is expensive. If you do it, put the threshold on a cosine of about 0.80 or above (a Spring AI score of 0.90) for this model, run it in shadow mode first (look up, but still call the model and compare), and sample the hits by hand. What this article did not test: any embedding model but one; streaming responses; multi-turn conversations; a larger cache than 30 distinct entries, where HNSW recall and search time begin to matter; Redis persistence and eviction; concurrent writers; real user traffic; and a hosted model’s latency. The tests show how a threshold trades hits against wrong answers for one model on one set of questions I wrote, not what your numbers will be.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.