Files

6 lines
340 B
Plaintext

# Where the milliseconds go on a lookup (110 questions, 30 entries cached, 2 vCPUs)
embed the question (MiniLM in-process) : median 2.4 ms, p95 3.6 ms
Redis KNN search with the tenant filter : median 0.3 ms, p95 2.5 ms (lookup time minus one embedding)
Timings drift run to run; the shape (embedding dominates, Redis is small) does not.