Add providers module: one app on the real OpenAI, Anthropic and Gemini Spring AI models against a local server in three wire formats; options, prompt caching, cost from price sheets, failover with retry layers measured
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,20 @@
|
||||
# One TicketService.raw() call, three providers, as the HTTP server saw it
|
||||
|
||||
OPENAI POST /v1/chat/completions
|
||||
top-level keys : messages, model
|
||||
model sent : gpt-6.1-sol
|
||||
system prompt : messages[0], role "system"
|
||||
sampling sent : nothing (the vendor's own defaults apply)
|
||||
|
||||
ANTHROPIC POST /v1/messages
|
||||
top-level keys : max_tokens, messages, model, system
|
||||
model sent : claude-sonnet-5-5
|
||||
system prompt : top-level "system" (a string)
|
||||
sampling sent : max_tokens=1024
|
||||
|
||||
GEMINI POST /v1beta/models/gemini-3.8-flash:generateContent
|
||||
top-level keys : contents, systemInstruction, generationConfig
|
||||
model sent : gemini-3.8-flash
|
||||
system prompt : top-level "systemInstruction".parts[0]
|
||||
sampling sent : temperature=0.7, topP=1.0
|
||||
|
||||
@@ -0,0 +1,20 @@
|
||||
# Options: portable, provider-specific, and the per-call trap
|
||||
|
||||
A. ChatClient.options(ChatOptions.builder().temperature(0.2).maxTokens(200)) on each provider
|
||||
OPENAI model=gpt-6.1-sol, max_tokens=200, temperature=0.2
|
||||
ANTHROPIC max_tokens=200, model=claude-sonnet-5-5, temperature=0.2
|
||||
GEMINI temperature=0.2, topP=1.0, maxOutputTokens=200
|
||||
|
||||
B. Provider-specific options through the same ChatClient call
|
||||
OPENAI reasoningEffort + promptCacheKey -> prompt_cache_key=tickets-v1, reasoning_effort=low
|
||||
ANTHROPIC topK(5) -> top_k=5
|
||||
GEMINI thinkingBudget(0) -> generationConfig {"temperature":0.7,"topP":1.0,"thinkingConfig":{"thinkingBudget":0}}
|
||||
|
||||
C. The trap: one built ChatOptions object handed to Prompt, then model.call(prompt)
|
||||
OPENAI ClassCastException (DefaultChatOptions cannot be cast to OpenAiChatOptions)
|
||||
ANTHROPIC no error; model sent "claude-haiku-4-5", max_tokens 4096, temperature absent
|
||||
GEMINI ClassCastException (DefaultChatOptions cannot be cast to GoogleGenAiChatOptions)
|
||||
|
||||
D. The other trap: OpenAI-specific options sent to the other two models
|
||||
ANTHROPIC no error; model sent "claude-haiku-4-5"
|
||||
GEMINI ClassCastException (OpenAiChatOptions cannot be cast to GoogleGenAiChatOptions)
|
||||
@@ -0,0 +1,18 @@
|
||||
# Prompt caching: what each client sends and what Usage reports
|
||||
|
||||
A. Anthropic: the strategy decides whether the system prompt carries a cache breakpoint
|
||||
strategy NONE system is a plain string
|
||||
call 1: promptTokens=6014 cacheRead=0 cacheWrite=0
|
||||
call 2: promptTokens=6014 cacheRead=0 cacheWrite=0
|
||||
strategy SYSTEM_ONLY system is a list of 1 block(s), cache_control on block 0: true
|
||||
call 1: promptTokens=3 cacheRead=0 cacheWrite=6011
|
||||
call 2: promptTokens=3 cacheRead=6011 cacheWrite=0
|
||||
|
||||
B. OpenAI and Gemini: nothing to switch on, but the order of the text decides the hit
|
||||
OPENAI timestamp AFTER the policy: call 1 cacheRead=0 call 2 cacheRead=6016 of 6019 prompt tokens
|
||||
OPENAI timestamp BEFORE the policy: call 1 cacheRead=0 call 2 cacheRead=0 of 6019 prompt tokens
|
||||
GEMINI timestamp AFTER the policy: call 1 cacheRead=0 call 2 cacheRead=6016 of 6019 prompt tokens
|
||||
GEMINI timestamp BEFORE the policy: call 1 cacheRead=0 call 2 cacheRead=0 of 6019 prompt tokens
|
||||
|
||||
C. A prompt below the vendor's minimum: the client still asks, the vendor does not cache
|
||||
Anthropic SYSTEM_ONLY, 847-character policy: cache_control sent=true, cacheRead=0 cacheWrite=0
|
||||
@@ -0,0 +1,15 @@
|
||||
# Cost of 100 ticket summaries with a ~6,000-token policy
|
||||
|
||||
Workload per request: system policy 24044 characters, ticket about 12 characters, answer 62 characters.
|
||||
Token counts are the fake server's estimate (characters / 4). Prices are per million tokens as read on 2026-10-09.
|
||||
|
||||
provider, model, price sheet cache miss cache hit saving
|
||||
OpenAI gpt-6.1-sol $1.22 $0.09 92.7%
|
||||
Anthropic claude-sonnet-5-5 $1.22 $0.09 92.5%
|
||||
Anthropic, clock inside the block $1.22 $1.52 -24.7%
|
||||
Gemini gemini-3.8-flash (to 2026) $0.46 $0.06 87.9%
|
||||
Gemini gemini-3.8-flash (2027) $0.92 $0.11 87.9%
|
||||
|
||||
"cache miss": OpenAI and Gemini with the changing clock text at the START of the system prompt; Anthropic with caching off.
|
||||
"cache hit": the clock text moved into the user message, so the system prompt never changes; Anthropic with SYSTEM_ONLY.
|
||||
The Anthropic "clock inside the block" row keeps the clock at the END of the one cached system block: it changes every request, so every request pays the cache write.
|
||||
@@ -0,0 +1,22 @@
|
||||
# Fallback: OpenAI first, then Anthropic, then Gemini
|
||||
|
||||
A. OpenAI fails with each status (no client retries). Does the call move on?
|
||||
OpenAI 400 -> thrown requests: openai=1 anthropic=0 gemini=0 trail: [openai: failed with status 400]
|
||||
OpenAI 401 -> thrown requests: openai=1 anthropic=0 gemini=0 trail: [openai: failed with status 401]
|
||||
OpenAI 429 -> answered requests: openai=1 anthropic=1 gemini=0 trail: [openai: failed with status 429, anthropic: ok]
|
||||
OpenAI 500 -> answered requests: openai=1 anthropic=1 gemini=0 trail: [openai: failed with status 500, anthropic: ok]
|
||||
OpenAI 503 -> answered requests: openai=1 anthropic=1 gemini=0 trail: [openai: failed with status 503, anthropic: ok]
|
||||
|
||||
B. The same 503 with the client's own retries switched on (maxRetries 2)
|
||||
OpenAI 503 -> answered requests: openai=3 anthropic=1 gemini=0
|
||||
|
||||
C. Gemini retries in two layers: the Google SDK (HttpRetryOptions.attempts) and Spring AI's RetryTemplate
|
||||
SDK attempts=1, RetryTemplate retries=0 -> 1 request(s) for one call
|
||||
SDK attempts=1, RetryTemplate retries=2 -> 3 request(s) for one call
|
||||
SDK attempts=3, RetryTemplate retries=0 -> 3 request(s) for one call
|
||||
SDK attempts=3, RetryTemplate retries=2 -> 9 request(s) for one call
|
||||
nothing configured at all -> 5 requests for one call, and it took more than 5 seconds: true
|
||||
|
||||
D. Everything down: the last provider's failure is the one you see
|
||||
all three 503 -> thrown requests: openai=1 anthropic=1 gemini=1
|
||||
trail: [openai: failed with status 503, anthropic: failed with status 503, gemini: failed with status 503]
|
||||
Reference in New Issue
Block a user