Add llm-gateway module: routing, failover below the tool-calling advisor, per-provider circuit breakers, dollar caps and token limits, tenant-keyed cache; real OpenAI and Anthropic models against a local fake
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,14 @@
|
||||
# Token rate limit: pre-consume, settle, refill
|
||||
|
||||
capacity 10000 tokens per hour
|
||||
take estimate 1500 -> available 8500
|
||||
settle: prompt 100 + completion 1000 = 1100 -> available 8900
|
||||
old rule (refund estimate minus PROMPT tokens only): would have refunded 1400 and left 100 charged
|
||||
|
||||
take 500, real total 2000 -> available 6900 (the shortfall is charged, not forgiven)
|
||||
|
||||
take 6900 -> available 0
|
||||
take 1000 -> rejected: Token rate limit reached for acme
|
||||
|
||||
30 minutes later -> available 5000 (greedy refill: 10000 per hour)
|
||||
60 more minutes -> available 10000 (capped at capacity)
|
||||
Reference in New Issue
Block a user