Add llm-gateway module: routing, failover below the tool-calling advisor, per-provider circuit breakers, dollar caps and token limits, tenant-keyed cache; real OpenAI and Anthropic models against a local fake

Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
Claude
2026-10-09 10:24:57 +00:00
parent 80cd21f89b
commit cdee85d3f4
51 changed files with 2404 additions and 0 deletions
+10
View File
@@ -0,0 +1,10 @@
# The gateway over HTTP (real beans, real models, fake vendors)
1. normal call -> 200 {"content":"30 days.","provider":"openai","model":"gpt-6.1-sol","promptTokens":47,"completionTokens":52,"costMicros":614,"modelCalls":1,"servedFromCache":false,"trail":["openai: ok"]}
2. openai returns 503 -> 200 {"content":"30 days.","provider":"anthropic","model":"claude-sonnet-5-5","promptTokens":47,"completionTokens":52,"costMicros":614,"modelCalls":1,"servedFromCache":false,"trail":["openai: failed with status 503","anthropic: ok"]}
3. both providers down -> 503 Retry-After=30 No provider could answer: [openai: failed with status 503, anthropic: failed with status 529]
4. unknown hint -> 400 Unknown model hint 'locla'. Known hints: [fast, smart]
5. tool not allow-listed -> 400 Tool 'drop_tables' is not on the gateway allow-list [refund_order]
6. tenant over its cap -> 429 Budget for poor would be exceeded: spent 0 of 1000 microdollars, this call may cost up to 2574
7. body claims tenant acme, header says poor -> 429 (the header wins)
8. no X-Tenant-Id header -> 400