Add llm-gateway module: routing, failover below the tool-calling advisor, per-provider circuit breakers, dollar caps and token limits, tenant-keyed cache; real OpenAI and Anthropic models against a local fake
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,19 @@
|
||||
# Where failover sits relative to the tool loop
|
||||
|
||||
A. failover below the advisor (this module)
|
||||
tool executions: 1
|
||||
trail: [openai: ok, openai: failed with status 503, claude: ok]
|
||||
answer: Refunded as refund-A17-1 (from claude)
|
||||
what claude was sent: UserMessage, AssistantMessage, ToolResponseMessage
|
||||
|
||||
B. retry above the advisor (one ChatClient per provider)
|
||||
tool executions: 2
|
||||
answer: Refunded as refund-A17-2
|
||||
|
||||
C. usage across a two-call tool loop (calls reported 120+15 and 160+12)
|
||||
usage on the final ChatResponse: prompt=280 completion=27
|
||||
usage summed per model call: prompt=280 completion=27
|
||||
|
||||
D. the same loop with the first call on a cheaper provider (0.75/3.75 then 2.00/10.00)
|
||||
priced per model call: 587 microdollars
|
||||
total usage at the last provider's price: 830 microdollars
|
||||
Reference in New Issue
Block a user