Files
spring-ai/semantic-cache/output/04-replay-600-requests.txt

17 lines
1.2 KiB
Plaintext

# 600 requests through the advisor, by threshold
Workload: 600 requests (seed 42), 30 intents asked with Zipf popularity in random phrasings, 15% off-topic.
121 distinct texts, so a plain exact-match cache could answer at most 479 of 600 (80%).
Cost model (ILLUSTRATIVE): $2.50 per million input tokens, $10.00 per million output tokens, tokens = characters / 4.
setup | model calls | hit rate | wrong | wrong of hits | model $ | saved
no cache | 600 | - | - | - | 0.4564 | -
cache, threshold 0.70 | 42 | 93.0% | 191 | 34.2% | 0.0316 | 93%
cache, threshold 0.80 | 62 | 89.7% | 7 | 1.3% | 0.0468 | 90%
cache, threshold 0.90 | 109 | 81.8% | 0 | 0.0% | 0.0827 | 82%
cache, threshold 0.95 | 119 | 80.2% | 0 | 0.0% | 0.0904 | 80%
cache, exact match only | 121 | 79.8% | 0 | 0.0% | 0.0919 | 80%
Average time to serve a cache hit (embed the question, search Redis, build the response): 4 ms at 0.70, 4 ms at 0.80, 3 ms at 0.90, 3 ms at 0.95
Saved dollars exclude the cost of embedding (local model here, CPU only) and of running Redis.