Add llm-gateway module: routing, failover below the tool-calling advisor, per-provider circuit breakers, dollar caps and token limits, tenant-keyed cache; real OpenAI and Anthropic models against a local fake
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,10 @@
|
||||
# Semantic cache mechanics (stand-in embeddings, threshold 0.92)
|
||||
|
||||
EmbeddingModel.embed(String) returns: float[]
|
||||
probe cosine hit at 0.92?
|
||||
same words, new order 0.985 true
|
||||
one word different (A18 for A17) 0.970 true
|
||||
different question, shares a few words 0.395 false
|
||||
|
||||
tenant globex asks the stored question: miss
|
||||
a cache keyed by feature only, tenant globex asks: 30 days, original packaging.
|
||||
Reference in New Issue
Block a user