Add llm-gateway module: routing, failover below the tool-calling advisor, per-provider circuit breakers, dollar caps and token limits, tenant-keyed cache; real OpenAI and Anthropic models against a local fake
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,11 @@
|
||||
# Routing: hints, per-target options, unknown hints
|
||||
|
||||
hint smart -> answered by claude (claude-model)
|
||||
model id sent to openai: openai-model
|
||||
model id sent to claude: claude-model
|
||||
maxTokens sent to both: 256 and 256
|
||||
|
||||
hint local -> answered by local; cloud calls made for it: 0
|
||||
|
||||
old registry, hint "locla" (typo for local): goes to claude, a cloud provider
|
||||
this Routes, hint "locla": Unknown model hint 'locla'. Known hints: [local, smart]
|
||||
@@ -0,0 +1,12 @@
|
||||
# Failover: which statuses move to the next provider
|
||||
|
||||
status second tried? outcome trail
|
||||
408 true answered [first: failed with status 408, second: ok]
|
||||
429 true answered [first: failed with status 429, second: ok]
|
||||
500 true answered [first: failed with status 500, second: ok]
|
||||
503 true answered [first: failed with status 503, second: ok]
|
||||
400 false rethrown [first: rejected with status 400, not retried]
|
||||
401 false rethrown [first: rejected with status 401, not retried]
|
||||
404 false rethrown [first: rejected with status 404, not retried]
|
||||
|
||||
both down: No provider could answer: [a: failed with status 503, b: failed with status 429]
|
||||
@@ -0,0 +1,11 @@
|
||||
# Circuit breakers: when they open and how they recover
|
||||
|
||||
old yml (window 10, threshold 50%): minimumNumberOfCalls is 100; the first request that skips openai is #11
|
||||
this module (window 10, minimum 5, threshold 50%): the first request that skips openai is #6
|
||||
|
||||
after 8 requests: openai breaker OPEN, claude breaker CLOSED
|
||||
openai was contacted 5 times of 8; claude answered 8 times
|
||||
|
||||
after the wait, one probe goes to openai (it has recovered): breaker CLOSED, openai calls 1
|
||||
|
||||
20 requests rejected with 400: breaker CLOSED, calls the breaker recorded: 0
|
||||
@@ -0,0 +1,19 @@
|
||||
# Where failover sits relative to the tool loop
|
||||
|
||||
A. failover below the advisor (this module)
|
||||
tool executions: 1
|
||||
trail: [openai: ok, openai: failed with status 503, claude: ok]
|
||||
answer: Refunded as refund-A17-1 (from claude)
|
||||
what claude was sent: UserMessage, AssistantMessage, ToolResponseMessage
|
||||
|
||||
B. retry above the advisor (one ChatClient per provider)
|
||||
tool executions: 2
|
||||
answer: Refunded as refund-A17-2
|
||||
|
||||
C. usage across a two-call tool loop (calls reported 120+15 and 160+12)
|
||||
usage on the final ChatResponse: prompt=280 completion=27
|
||||
usage summed per model call: prompt=280 completion=27
|
||||
|
||||
D. the same loop with the first call on a cheaper provider (0.75/3.75 then 2.00/10.00)
|
||||
priced per model call: 587 microdollars
|
||||
total usage at the last provider's price: 830 microdollars
|
||||
@@ -0,0 +1,8 @@
|
||||
# Real OpenAI and Anthropic models: failover in the middle of a tool loop
|
||||
|
||||
trail: [openai: ok, openai: failed with status 503, anthropic: ok]
|
||||
answered by anthropic; tool executions 1; requests: openai 2, anthropic 1
|
||||
|
||||
request 1 to OpenAI: model=gpt-6.1-sol, max_tokens=256, tools=[refund_order]
|
||||
request to Anthropic: model=claude-sonnet-5-5 max_tokens=256
|
||||
conversation it received: [assistant:tool_use(id=call_1, name=refund_order), user:tool_result(tool_use_id=call_1)]
|
||||
@@ -0,0 +1,9 @@
|
||||
# What the vendor SDKs throw, and what the gateway does with it
|
||||
|
||||
vendor status exception transient? HTTP requests for one call
|
||||
openai 400 com.openai.errors.BadRequestException false 1
|
||||
openai 401 com.openai.errors.UnauthorizedException false 1
|
||||
openai 429 com.openai.errors.RateLimitException true 1
|
||||
openai 503 com.openai.errors.InternalServerException true 1
|
||||
anthropic 400 com.anthropic.errors.BadRequestException false 1
|
||||
anthropic 529 com.anthropic.errors.InternalServerException true 1
|
||||
@@ -0,0 +1,11 @@
|
||||
# Dollar caps: reserve, settle, and the limits of an estimate
|
||||
|
||||
cap 5000 microdollars; every call really costs 2000; maxTokens 256 (estimate about 2,570)
|
||||
call 1: answered, cost 2000, spent so far 2000
|
||||
call 2: answered, cost 2000, spent so far 4000
|
||||
call 3: rejected before the provider was called (provider calls so far: 2)
|
||||
|
||||
cap 8000; prompt estimated at 100 tokens but the provider counts 3000 (code, other scripts, images do this)
|
||||
spent after 2 calls: 12200 (cap 8000); the second call was admitted on its estimate
|
||||
|
||||
cap 2700, a two-call tool loop: rejected on the SECOND model call; tool executions so far: 1
|
||||
@@ -0,0 +1,5 @@
|
||||
# 64 simultaneous requests against a cap that fits 10
|
||||
|
||||
cap 10000, each call 1000, 64 threads at once
|
||||
check then record: admitted 64, spent 64000
|
||||
reserve then settle: admitted 10
|
||||
@@ -0,0 +1,14 @@
|
||||
# Token rate limit: pre-consume, settle, refill
|
||||
|
||||
capacity 10000 tokens per hour
|
||||
take estimate 1500 -> available 8500
|
||||
settle: prompt 100 + completion 1000 = 1100 -> available 8900
|
||||
old rule (refund estimate minus PROMPT tokens only): would have refunded 1400 and left 100 charged
|
||||
|
||||
take 500, real total 2000 -> available 6900 (the shortfall is charged, not forgiven)
|
||||
|
||||
take 6900 -> available 0
|
||||
take 1000 -> rejected: Token rate limit reached for acme
|
||||
|
||||
30 minutes later -> available 5000 (greedy refill: 10000 per hour)
|
||||
60 more minutes -> available 10000 (capped at capacity)
|
||||
@@ -0,0 +1,10 @@
|
||||
# Semantic cache mechanics (stand-in embeddings, threshold 0.92)
|
||||
|
||||
EmbeddingModel.embed(String) returns: float[]
|
||||
probe cosine hit at 0.92?
|
||||
same words, new order 0.985 true
|
||||
one word different (A18 for A17) 0.970 true
|
||||
different question, shares a few words 0.395 false
|
||||
|
||||
tenant globex asks the stored question: miss
|
||||
a cache keyed by feature only, tenant globex asks: 30 days, original packaging.
|
||||
@@ -0,0 +1,10 @@
|
||||
# The gateway over HTTP (real beans, real models, fake vendors)
|
||||
|
||||
1. normal call -> 200 {"content":"30 days.","provider":"openai","model":"gpt-6.1-sol","promptTokens":47,"completionTokens":52,"costMicros":614,"modelCalls":1,"servedFromCache":false,"trail":["openai: ok"]}
|
||||
2. openai returns 503 -> 200 {"content":"30 days.","provider":"anthropic","model":"claude-sonnet-5-5","promptTokens":47,"completionTokens":52,"costMicros":614,"modelCalls":1,"servedFromCache":false,"trail":["openai: failed with status 503","anthropic: ok"]}
|
||||
3. both providers down -> 503 Retry-After=30 No provider could answer: [openai: failed with status 503, anthropic: failed with status 529]
|
||||
4. unknown hint -> 400 Unknown model hint 'locla'. Known hints: [fast, smart]
|
||||
5. tool not allow-listed -> 400 Tool 'drop_tables' is not on the gateway allow-list [refund_order]
|
||||
6. tenant over its cap -> 429 Budget for poor would be exceeded: spent 0 of 1000 microdollars, this call may cost up to 2574
|
||||
7. body claims tenant acme, header says poor -> 429 (the header wins)
|
||||
8. no X-Tenant-Id header -> 400
|
||||
Reference in New Issue
Block a user