Add providers module: one app on the real OpenAI, Anthropic and Gemini Spring AI models against a local server in three wire formats; options, prompt caching, cost from price sheets, failover with retry layers measured
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,15 @@
|
||||
# Cost of 100 ticket summaries with a ~6,000-token policy
|
||||
|
||||
Workload per request: system policy 24044 characters, ticket about 12 characters, answer 62 characters.
|
||||
Token counts are the fake server's estimate (characters / 4). Prices are per million tokens as read on 2026-10-09.
|
||||
|
||||
provider, model, price sheet cache miss cache hit saving
|
||||
OpenAI gpt-6.1-sol $1.22 $0.09 92.7%
|
||||
Anthropic claude-sonnet-5-5 $1.22 $0.09 92.5%
|
||||
Anthropic, clock inside the block $1.22 $1.52 -24.7%
|
||||
Gemini gemini-3.8-flash (to 2026) $0.46 $0.06 87.9%
|
||||
Gemini gemini-3.8-flash (2027) $0.92 $0.11 87.9%
|
||||
|
||||
"cache miss": OpenAI and Gemini with the changing clock text at the START of the system prompt; Anthropic with caching off.
|
||||
"cache hit": the clock text moved into the user message, so the system prompt never changes; Anthropic with SYSTEM_ONLY.
|
||||
The Anthropic "clock inside the block" row keeps the clock at the END of the one cached system block: it changes every request, so every request pays the cache write.
|
||||
Reference in New Issue
Block a user