# Where the milliseconds go on a lookup (110 questions, 30 entries cached, 2 vCPUs) embed the question (MiniLM in-process) : median 2.4 ms, p95 3.6 ms Redis KNN search with the tenant filter : median 0.3 ms, p95 2.5 ms (lookup time minus one embedding) Timings drift run to run; the shape (embedding dominates, Redis is small) does not.