Add llm-gateway module: routing, failover below the tool-calling advisor, per-provider circuit breakers, dollar caps and token limits, tenant-keyed cache; real OpenAI and Anthropic models against a local fake

Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
Claude
2026-10-09 10:24:57 +00:00
parent 80cd21f89b
commit cdee85d3f4
51 changed files with 2404 additions and 0 deletions
+14
View File
@@ -0,0 +1,14 @@
# Token rate limit: pre-consume, settle, refill
capacity 10000 tokens per hour
take estimate 1500 -> available 8500
settle: prompt 100 + completion 1000 = 1100 -> available 8900
old rule (refund estimate minus PROMPT tokens only): would have refunded 1400 and left 100 charged
take 500, real total 2000 -> available 6900 (the shortfall is charged, not forgiven)
take 6900 -> available 0
take 1000 -> rejected: Token rate limit reached for acme
30 minutes later -> available 5000 (greedy refill: 10000 per hour)
60 more minutes -> available 10000 (capped at capacity)