Add providers module: one app on the real OpenAI, Anthropic and Gemini Spring AI models against a local server in three wire formats; options, prompt caching, cost from price sheets, failover with retry layers measured
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,18 @@
|
||||
# Prompt caching: what each client sends and what Usage reports
|
||||
|
||||
A. Anthropic: the strategy decides whether the system prompt carries a cache breakpoint
|
||||
strategy NONE system is a plain string
|
||||
call 1: promptTokens=6014 cacheRead=0 cacheWrite=0
|
||||
call 2: promptTokens=6014 cacheRead=0 cacheWrite=0
|
||||
strategy SYSTEM_ONLY system is a list of 1 block(s), cache_control on block 0: true
|
||||
call 1: promptTokens=3 cacheRead=0 cacheWrite=6011
|
||||
call 2: promptTokens=3 cacheRead=6011 cacheWrite=0
|
||||
|
||||
B. OpenAI and Gemini: nothing to switch on, but the order of the text decides the hit
|
||||
OPENAI timestamp AFTER the policy: call 1 cacheRead=0 call 2 cacheRead=6016 of 6019 prompt tokens
|
||||
OPENAI timestamp BEFORE the policy: call 1 cacheRead=0 call 2 cacheRead=0 of 6019 prompt tokens
|
||||
GEMINI timestamp AFTER the policy: call 1 cacheRead=0 call 2 cacheRead=6016 of 6019 prompt tokens
|
||||
GEMINI timestamp BEFORE the policy: call 1 cacheRead=0 call 2 cacheRead=0 of 6019 prompt tokens
|
||||
|
||||
C. A prompt below the vendor's minimum: the client still asks, the vendor does not cache
|
||||
Anthropic SYSTEM_ONLY, 847-character policy: cache_control sent=true, cacheRead=0 cacheWrite=0
|
||||
Reference in New Issue
Block a user