Add llm-gateway module: routing, failover below the tool-calling advisor, per-provider circuit breakers, dollar caps and token limits, tenant-keyed cache; real OpenAI and Anthropic models against a local fake
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,5 @@
|
||||
# 64 simultaneous requests against a cap that fits 10
|
||||
|
||||
cap 10000, each call 1000, 64 threads at once
|
||||
check then record: admitted 64, spent 64000
|
||||
reserve then settle: admitted 10
|
||||
Reference in New Issue
Block a user