Add llm-gateway module: routing, failover below the tool-calling advisor, per-provider circuit breakers, dollar caps and token limits, tenant-keyed cache; real OpenAI and Anthropic models against a local fake

Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
Claude
2026-10-09 10:24:57 +00:00
parent 80cd21f89b
commit cdee85d3f4
51 changed files with 2404 additions and 0 deletions
+11
View File
@@ -0,0 +1,11 @@
# Dollar caps: reserve, settle, and the limits of an estimate
cap 5000 microdollars; every call really costs 2000; maxTokens 256 (estimate about 2,570)
call 1: answered, cost 2000, spent so far 2000
call 2: answered, cost 2000, spent so far 4000
call 3: rejected before the provider was called (provider calls so far: 2)
cap 8000; prompt estimated at 100 tokens but the provider counts 3000 (code, other scripts, images do this)
spent after 2 calls: 12200 (cap 8000); the second call was admitted on its estimate
cap 2700, a two-call tool loop: rejected on the SECOND model call; tool executions so far: 1