Skip to main content

Anthropic Claude vs OpenAI vs Gemini in Spring AI 2.0: Switching Providers and Comparing Cost

You built a feature on one AI provider. Then someone asks the questions every team asks sooner or later: could we try a different one, what would it cost, and what happens when the one we use has a bad afternoon? In Spring AI the answer to the first question is “mostly yes, it is a configuration change”. The word mostly is where the surprises live. This article takes one small application and runs it on OpenAI, Anthropic Claude and Google Gemini through the three real Spring AI model classes. It looks at what each one sends over the wire, which options carry over and which silently do not, how prompt caching differs, what a batch of requests costs on each price sheet, and how to fail over from one provider to the next without multiplying your retries by accident. The depth is in expandable sections, so you can read straight through or open only what you need.
Versions, and an honest limit. Spring Boot 4.1.1, Spring AI 2.0.1 and Java 25. The code is the providers module of asmhatre/spring-ai, and every console block below is quoted from a file under its output/ directory. I did not call any vendor’s API; there are no keys in my build environment. The real OpenAiChatModel, AnthropicChatModel and GoogleGenAiChatModel talk to a local server that answers in the three wire formats, so what each client sends and how it reacts to a reply or an error is real. Three things are simulated, and the article says so wherever they matter: token counts are characters divided by four, the answer is a fixed string, and cache hits follow the rules the vendors document (thresholds, prefix matching), applied to the request that actually arrived. So the cost table is a worked example of the published price sheets, not a bill, and I did not measure latency or answer quality at all. Prices and cache rules are as I read them on 9 October 2026; they change often.

Switching providers is one class, if you let it be

Spring AI gives every provider the same interface, ChatModel. Your application code asks a ChatModel a question and gets an answer; it does not need to know whether the thing behind it is OpenAI, Claude or Gemini. The only code that knows is the code that builds the model. If you keep that in one place, switching providers is a one-line change there, and nowhere else.
TicketServiceyour codeChatModelthe interfaceOpenAiChatModelchat/completionsAnthropicChatModelmessagesGoogleGenAiChatModelgenerateContentOpenAIAnthropicGoogle
Read the picture left to right. Everything to the left of the interface is your application and never changes. Everything to the right is replaceable, and the middle column is the only part that is specific to a provider. The application here summarises a support ticket into a small record; it holds a ChatModel and nothing else. This is TicketService.java; it has no import from any vendor package.
public class TicketService {

    private final ChatClient client;

    private final String policy;

    public TicketService(ChatModel model, String policy) {
        this.client = ChatClient.create(model);
        this.policy = policy;
    }

    public Summary summarise(String ticket) {
        return client.prompt().system(policy).user(ticket).call().entity(Summary.class);
    }

    public String raw(String ticket) {
        return client.prompt().system(policy).user(ticket).call().content();
    }
}
And this is the one place that knows the vendors, ProviderModels.java. Each branch builds the real model class with its API key, base URL, model id and retry count. In my tests the base URL points at the local server; in your app you would leave it at its default and read the key from configuration.
    public static ChatModel create(Provider p, Settings s, String model) {
        return switch (p) {
            case OPENAI -> OpenAiChatModel.builder()
                    .options(OpenAiChatOptions.builder().baseUrl(s.baseUrl() + "/v1").apiKey(s.apiKey())
                            .model(model).maxRetries(s.maxRetries()).build())
                    .build();
            case ANTHROPIC -> AnthropicChatModel.builder()
                    .options(AnthropicChatOptions.builder().baseUrl(s.baseUrl()).apiKey(s.apiKey())
                            .model(model).maxTokens(1024).maxRetries(s.maxRetries())
                            .cacheOptions(AnthropicCacheOptions.builder().strategy(s.anthropicCache()).build())
                            .build())
                    .build();
            case GEMINI -> GoogleGenAiChatModel.builder()
                    .genAiClient(Client.builder().apiKey(s.apiKey())
                            .httpOptions(HttpOptions.builder().baseUrl(s.baseUrl())
                            .retryOptions(HttpRetryOptions.builder().attempts(s.maxRetries() + 1).initialDelay(0.001).build()).build()).build())
                    .options(GoogleGenAiChatOptions.builder().model(model).build())
                    .retryTemplate(new RetryTemplate(RetryPolicy.builder().maxRetries(0)
                            .delay(Duration.ofMillis(1)).build()))
                    .build();
        };
    }
Two details in that code are worth a second look. Gemini is configured through Google’s own client object (Client), not through Spring properties, and it has two places to set retries, which the failover section comes back to. And the Gemini auto-configuration I looked at has no base-URL property at all (its connection properties are the API key, project, location, credentials and a Vertex AI switch), which is why this module builds the models by hand: I could not point the starter at a local server with properties alone.
Going deeper: switching with properties instead of code
With the Spring AI starters you normally do not write a factory. You add the starter for each provider you might use and pick one with the spring.ai.model.chat property (for example openai, anthropic or google-genai); only the selected provider’s ChatModel bean is created. The ChatClient article has a test that proves the switch. I did not repeat it for all three providers here, so treat the property names as something to confirm against the reference for your version. The factory in this module does the same job in plain code, which is easier to point at a test server and easier to read when you want two providers live at once, as failover needs.

The same call becomes three different requests

Because the application code is identical, it is easy to assume the three requests are too. They are not. I made the same call (the same system prompt, the same ticket) through each model and recorded what the server received (01-wire.txt, from WireTest.java):
OPENAI    POST /v1/chat/completions
          top-level keys : messages, model
          model sent     : gpt-6.1-sol
          system prompt  : messages[0], role "system"
          sampling sent  : nothing (the vendor's own defaults apply)

ANTHROPIC POST /v1/messages
          top-level keys : max_tokens, messages, model, system
          model sent     : claude-sonnet-5-5
          system prompt  : top-level "system" (a string)
          sampling sent  : max_tokens=1024

GEMINI    POST /v1beta/models/gemini-3.8-flash:generateContent
          top-level keys : contents, systemInstruction, generationConfig
          model sent     : gemini-3.8-flash
          system prompt  : top-level "systemInstruction".parts[0]
          sampling sent  : temperature=0.7, topP=1.0
Three differences stand out. First, where the system prompt goes: OpenAI makes it the first message with the role “system”, Anthropic has a separate top-level system field, and Gemini calls it systemInstruction. Spring AI hides that, which is the point of the interface. Second, what is required: Anthropic’s API requires max_tokens, so every request carries one (1024 in my configuration). OpenAI and Gemini do not need it. Third, and easiest to miss, what is sent when you set nothing: OpenAI and Anthropic send no sampling settings and the vendor’s own defaults apply, but the Gemini model sends temperature=0.7 and topP=1.0 on every request.
An option you never set can still differ. If you switch from OpenAI to Gemini and the answers feel more varied, check the request before blaming the model: Spring AI’s Gemini model fills in a temperature of 0.7 when you have not chosen one. Set temperature explicitly on every provider if you care about consistent behaviour across a switch.
Going deeper: what else I compared, and what I did not
I compared only the parts of the request that every provider has. The three APIs also differ in how they represent tools, images and structured output, which this article leaves to the posts that cover them: tool calling, images across three wire formats and structured output. The response side differs as well (OpenAI wraps the answer in choices, Anthropic in a list of content blocks, Gemini in candidates), but Spring AI maps all three into the same ChatResponse, so you only see it if you read the raw HTTP.

Options: portable until you hold them the wrong way

Every model has settings beyond the prompt: how random the answer should be, how long it may be, and a long tail of vendor-specific switches. Spring AI has a portable layer, ChatOptions, with the settings every provider has (temperature, maximum tokens, top-p), and a provider-specific class for the rest (OpenAiChatOptions, AnthropicChatOptions, GoogleGenAiChatOptions). Used through ChatClient, the portable layer does what you hope. This is the same .options(...) call on each provider (02-options.txt, from OptionsTest.java):
A. ChatClient.options(ChatOptions.builder().temperature(0.2).maxTokens(200)) on each provider
  OPENAI    model=gpt-6.1-sol, max_tokens=200, temperature=0.2
  ANTHROPIC max_tokens=200, model=claude-sonnet-5-5, temperature=0.2
  GEMINI    temperature=0.2, topP=1.0, maxOutputTokens=200

B. Provider-specific options through the same ChatClient call
  OPENAI    reasoningEffort + promptCacheKey -> prompt_cache_key=tickets-v1, reasoning_effort=low
  ANTHROPIC topK(5)                          -> top_k=5
  GEMINI    thinkingBudget(0)                -> generationConfig {"temperature":0.7,"topP":1.0,"thinkingConfig":{"thinkingBudget":0}}
The first block shows one portable call, temperature(0.2).maxTokens(200), landing in each vendor’s own field names: OpenAI sends max_tokens, Anthropic max_tokens too, and Gemini maxOutputTokens inside a generationConfig object. The model id configured when the model was built survives in all three. The second block shows provider-specific settings (OpenAI’s reasoning effort and cache key, Anthropic’s top-k, Gemini’s thinking budget) going through the same call, each appearing in its own place. Now the part that bit me. There is another way to give options to a model: build a finished options object and put it in a Prompt, then call the model directly. It looks equivalent and it is not (second half of 02-options.txt):
C. The trap: one built ChatOptions object handed to Prompt, then model.call(prompt)
  OPENAI    ClassCastException (DefaultChatOptions cannot be cast to OpenAiChatOptions)
  ANTHROPIC no error; model sent "claude-haiku-4-5", max_tokens 4096, temperature absent
  GEMINI    ClassCastException (DefaultChatOptions cannot be cast to GoogleGenAiChatOptions)

D. The other trap: OpenAI-specific options sent to the other two models
  ANTHROPIC no error; model sent "claude-haiku-4-5"
  GEMINI    ClassCastException (OpenAiChatOptions cannot be cast to GoogleGenAiChatOptions)
A plain ChatOptions built that way crashes the OpenAI and Gemini models with a ClassCastException, because they expect their own options class. Anthropic does something quieter and worse: it accepts the object without complaint, ignores the temperature and token limit you set, and sends claude-haiku-4-5 with max_tokens 4096, not the claude-sonnet-5-5 and 1024 the model was configured with. The OpenAI-specific options sent to Anthropic did the same model switch. In the output above, nothing failed; only the request on the wire shows it.
Pass options through ChatClient.options(...) with a builder, not as a built object inside a Prompt. The ChatClient path merges your settings into the model’s defaults; the direct path can replace them, and with Anthropic it replaced the model id. If you must call a ChatModel directly, use that provider’s own options class, never another provider’s and never the bare ChatOptions.
Going deeper: why I trust the wire and not the types
All three behaviours above come from capturing the HTTP request, not from reading source, so they describe Spring AI 2.0.1 with these exact client libraries and may change in a later release. The lesson that will survive a version bump is the method: when you switch providers, log or capture one real request from each and compare. The observability article puts the model name on its metrics, which would catch the Anthropic model switch in production.

Prompt caching: three providers, three behaviours

Many applications send the same long instructions with every request: a policy, a persona, a set of examples. You pay for those input tokens every time. Prompt caching lets the provider remember the beginning of a prompt it has seen recently, and charge a fraction for the repeated part. All three vendors offer it, but you turn it on in three different ways, and one rule is shared by all: the cache matches on the start of the prompt. Anything that changes early ruins everything after it.
Request 1 and request 2, changing text last: the policy is reusedlong policy, identical both times (cached, cheap)clock + ticket (full price)Request 1 and request 2, changing text first: nothing is reusedclocklong policy, but it now follows text that differs, so it cannot match (full price)The cache compares from the first character; the first difference ends the match.
Both rows contain the same words. In the top row the unchanging policy leads, so a second request shares a long prefix with the first and the provider can reuse that work. In the bottom row a clock stamp leads, so the very first characters already differ and the policy behind it is billed in full every time, even though it never changed. Here is that, measured through Spring AI’s own usage numbers, for OpenAI and Gemini, which need nothing switched on (03-caching.txt, from CachingTest.java). The prompt is about 6,000 tokens:
B. OpenAI and Gemini: nothing to switch on, but the order of the text decides the hit
  OPENAI  timestamp AFTER the policy:        call 1 cacheRead=0     call 2 cacheRead=6016 of 6019 prompt tokens
  OPENAI  timestamp BEFORE the policy:       call 1 cacheRead=0     call 2 cacheRead=0 of 6019 prompt tokens
  GEMINI  timestamp AFTER the policy:        call 1 cacheRead=0     call 2 cacheRead=6016 of 6019 prompt tokens
  GEMINI  timestamp BEFORE the policy:       call 1 cacheRead=0     call 2 cacheRead=0 of 6019 prompt tokens
With the clock after the policy, the second request reads 6,016 of its 6,019 prompt tokens from cache. With the clock before it, zero. Remember what produced those numbers: my local server applying the documented rule (a shared prefix over a minimum length is reused) to the request it received. The part that is genuinely measured is the request: it is identical in both layouts except for where the clock sits, and Spring AI’s Usage object reports the cached count under one portable name, getCacheReadInputTokens(), for both providers. Anthropic is different: you must ask. In Spring AI that is a cache strategy on the model’s options, set once when the model is built (ProviderModels.java):
                            .cacheOptions(AnthropicCacheOptions.builder().strategy(s.anthropicCache()).build())
                            .build())
With SYSTEM_ONLY the client changes the shape of the request: the system prompt stops being a plain string and becomes a list of blocks, the last of which carries a cache_control marker meaning “cache everything up to here”. The usage numbers then show the two sides of Anthropic’s pricing, a write the first time and a read after (first half of 03-caching.txt):
A. Anthropic: the strategy decides whether the system prompt carries a cache breakpoint
  strategy NONE         system is a plain string
    call 1: promptTokens=6014 cacheRead=0 cacheWrite=0
    call 2: promptTokens=6014 cacheRead=0 cacheWrite=0
  strategy SYSTEM_ONLY  system is a list of 1 block(s), cache_control on block 0: true
    call 1: promptTokens=3 cacheRead=0 cacheWrite=6011
    call 2: promptTokens=3 cacheRead=6011 cacheWrite=0
Without the strategy, both calls report 6,014 prompt tokens and no cache activity. With it, the first call reports 6,011 tokens written to the cache and only 3 ordinary prompt tokens; the second reads those 6,011 back. Note how the number called “prompt tokens” changed from 6,014 to 3 for the same text: the next section explains why that matters for cost.
For Anthropic, the cached block is the whole system block. If you put the changing text (a timestamp, a user name) inside the same system string as the long policy, every request has a different block, so every request writes a new cache entry and never reads one. Keep the changing parts in the user message, or in a separate block after the cached one. The cost table below includes this mistake on purpose.
Going deeper: the documented rules I simulated
These are the rules the local server applies, as I read them from the vendors’ documentation on 9 October 2026; the first line of each is the one that decides whether your prompt is eligible at all.
  • Anthropic: a minimum of 512 tokens for Claude Sonnet 5.5 (other models range from 512 to 4,096); entries last five minutes by default and refresh on use, with a one-hour option at a higher write price; up to four breakpoints per request; the cached prefix is tools, then system, then messages, up to the marked block; usage reports cache_creation_input_tokens and cache_read_input_tokens. Spring AI exposes the strategies NONE, TOOLS_ONLY, SYSTEM_ONLY, SYSTEM_AND_TOOLS and CONVERSATION_HISTORY and the lifetimes FIVE_MINUTES and ONE_HOUR.
  • OpenAI: automatic, with a minimum of 1,024 visible input tokens on the newest models; “the entire rendered prefix” must match; a prompt_cache_key option exists to keep accounting separate per customer or workspace.
  • Gemini: implicit caching is on by default for Gemini 2.5 and newer; the minimum for Gemini 3.8 Flash is 4,096 tokens. There is also an explicit cached-content feature in Spring AI’s Gemini options (cachedContentName, autoCacheThreshold) that I did not run.
I also checked the case where the prompt is too short. With an 847-character policy and SYSTEM_ONLY, Spring AI still sent the cache marker; the vendor would simply not cache it, and the usage showed no read and no write (03-caching.txt, last line). That fails silently, so look at the usage numbers once rather than assume.

What it costs: a worked example, with one trap

To turn token counts into money you need two things: the numbers a response reports, and a price per million tokens for each kind. Spring AI gives you the first through Usage. The catch is that getPromptTokens() does not mean the same thing for every provider. For OpenAI and Gemini it is the total input, with the cached part counted inside it. For Anthropic it is only the tokens after the cache breakpoint; cached tokens are reported separately. Add them up the same way for all three and you will be wrong for at least one. This is the small class that handles it, Cost.java. For OpenAI and Gemini it subtracts the cache reads from the prompt tokens to find the part billed at the normal rate; for Anthropic it does not:
    public static Billed of(Provider provider, Usage u, Prices p) {
        long prompt = u.getPromptTokens();
        long read = orZero(u.getCacheReadInputTokens());
        long write = orZero(u.getCacheWriteInputTokens());
        long uncached = switch (provider) {
            case ANTHROPIC -> prompt;
            case OPENAI, GEMINI -> prompt - read;
        };
        long out = u.getCompletionTokens();
        BigDecimal usd = p.input().multiply(BigDecimal.valueOf(uncached))
                .add(p.cacheRead().multiply(BigDecimal.valueOf(read)))
                .add(p.cacheWrite().multiply(BigDecimal.valueOf(write)))
                .add(p.output().multiply(BigDecimal.valueOf(out)))
                .divide(MILLION, 8, RoundingMode.HALF_UP);
        return new Billed(uncached, read, write, out, usd);
    }
Then I ran 100 requests per setup, each with the same 24,044-character policy (about 6,000 tokens) and a short ticket, and priced the usage with each vendor’s published rates (04-cost.txt, from CostTest.java):
provider, model, price sheet         cache miss    cache hit    saving
OpenAI gpt-6.1-sol                        $1.22        $0.09     92.7%
Anthropic claude-sonnet-5-5               $1.22        $0.09     92.5%
Anthropic, clock inside the block         $1.22        $1.52    -24.7%
Gemini gemini-3.8-flash (to 2026)         $0.46        $0.06     87.9%
Gemini gemini-3.8-flash (2027)            $0.92        $0.11     87.9%
Read the first column as “caching is not helping” and the second as “caching is working”. For OpenAI and Gemini that means the changing text is at the start versus the end; for Anthropic it means the strategy is off versus on. When caching works, the bill for the repeated policy nearly disappears: about 93% less for OpenAI and Anthropic and about 88% less for Gemini on these price sheets. The row in the middle is the trap from the previous section: Anthropic with caching switched on but the clock inside the cached block costs more than not caching at all, because every request pays the higher cache-write price and never earns a read.
OpenAI, miss$1.22OpenAI, hit$0.09Claude, off$1.22Claude, hit$0.09Claude, clock in block$1.52Gemini, miss$0.46Gemini, hit$0.06cost of 100 requests, ~6,000-token policy (lower is better)
The chart is the same table as bars. The red bars are the same $1.22 for OpenAI and Claude because their input and output prices happen to be identical on these sheets, and Gemini’s Flash model is cheaper per token. Do not read a ranking into this: the models are different sizes and the numbers say nothing about quality, which I did not test.
Going deeper: the price sheets I used, and what is missing from the sum
Prices are dollars per million tokens, standard tier, short context, read on 9 October 2026, and they are written into CostTest.java with their sources.
Provider and modelInputCached input (read)Cache writeOutput
OpenAI gpt-6.1-sol$2.00$0.10not priced here$10.00
Anthropic claude-sonnet-5-5$2.00$0.10$2.50 (5 minutes)$10.00
Gemini gemini-3.8-flash, through 31 Dec 2026$0.75$0.075storage $0.50 per million tokens per hour (explicit caching only)$3.75
Gemini gemini-3.8-flash, from 1 Jan 2027$1.50$0.15storage $1.00 per million tokens per hour$7.50
The Gemini row for 2027 is in the table because the page lists a scheduled doubling, and the test confirms that the same usage costs exactly twice as much on the new sheet. What the sum leaves out: reasoning or “thinking” tokens (Gemini bills them as output, and other models have similar settings), images, tool-call overhead, batch discounts, long-context surcharges and the one-hour cache lifetime. My answers are 16 tokens long, so output hardly shows; a model that writes long answers would shift the picture toward the output price.
Going deeper: what I did not measure, and how to
The plan for this article included a latency comparison. I did not produce one, deliberately. My server answers instantly, so any latency it showed would be the speed of a local test server, and a number like that in a table looks like data when it is not. Real latency depends on the model, the prompt length, the region, the time of day and whether the answer is streamed (time to first token and time to last token are different questions). To measure it for your own workload, record the duration of each call in production, tagged with provider and model, using the metrics Spring AI already emits; the observability article shows how.

When a provider fails: fail over, but count your retries

Providers have bad afternoons: overloaded responses, rate limits, brief outages. If you have the same feature set up on two providers, you can try the second when the first fails. That is failover, and it is a small amount of code. The harder part is deciding which failures deserve it, and knowing how many requests one failure really causes. Here is the heart of FallbackChatModel.java. It is a ChatModel itself, so the rest of the application does not know it is there. It tries each provider in turn, writes down what happened, and moves on only if a rule you give it says the failure is worth moving on from:
    public synchronized ChatResponse call(Prompt prompt) {
        trail.clear();
        RuntimeException last = null;
        for (Target t : targets) {
            try {
                ChatResponse r = t.model().call(prompt);
                trail.add(t.name() + ": ok");
                return r;
            }
            catch (RuntimeException e) {
                trail.add(t.name() + ": failed with status " + Failures.status(e));
                last = e;
                if (!shouldFallBack.test(e)) {
                    throw e;
                }
            }
        }
        throw last;
    }
The rule is Failures.java. The three vendors’ libraries throw three unrelated families of exception, and Gemini’s arrives wrapped inside a plain RuntimeException, so the code walks the chain of causes to find an HTTP status. Overload (429), timeouts (408) and server errors (5xx) are worth a second provider. A 400 (your request is malformed) or a 401 (your key is wrong) are not: another provider would only hide a bug you need to see. Here is each case, with OpenAI as the first choice (05-fallback.txt, from FallbackTest.java):
A. OpenAI fails with each status (no client retries). Does the call move on?
  OpenAI 400 -> thrown   requests: openai=1 anthropic=0 gemini=0  trail: [openai: failed with status 400]
  OpenAI 401 -> thrown   requests: openai=1 anthropic=0 gemini=0  trail: [openai: failed with status 401]
  OpenAI 429 -> answered requests: openai=1 anthropic=1 gemini=0  trail: [openai: failed with status 429, anthropic: ok]
  OpenAI 500 -> answered requests: openai=1 anthropic=1 gemini=0  trail: [openai: failed with status 500, anthropic: ok]
  OpenAI 503 -> answered requests: openai=1 anthropic=1 gemini=0  trail: [openai: failed with status 503, anthropic: ok]
A 400 and a 401 stop at OpenAI with one request and nothing sent to Anthropic. A 429, 500 or 503 moves on, and the trail records the sequence. This is the table I would want in an incident review, which is why the class keeps the trail. Now the second half of the question, how many requests one failure causes. Each provider’s client library retries on its own before your fallback ever sees the error. With OpenAI’s retry count set to 2, a single 503 became three requests to OpenAI before the one to Anthropic (05-fallback.txt, section B):
B. The same 503 with the client's own retries switched on (maxRetries 2)
  OpenAI 503 -> answered requests: openai=3 anthropic=1 gemini=0
Gemini is the surprising one. It retries in two layers: the Google client library has its own retry setting, and Spring AI’s model wraps the call in a Spring RetryTemplate as well. The two multiply (05-fallback.txt, section C):
Your callone failureRetryTemplate3 attemptsGoogle client3 attempts each9 requests3 attempts x 3 attempts = 9 HTTP requests for one call that never succeeded
The figure is the arithmetic from the next transcript. A Spring retry of three attempts wraps a client library that also makes three, so one doomed call makes nine requests. Here are the combinations, and what happens if you configure nothing (05-fallback.txt):
C. Gemini retries in two layers: the Google SDK (HttpRetryOptions.attempts) and Spring AI's RetryTemplate
  SDK attempts=1, RetryTemplate retries=0 -> 1 request(s) for one call
  SDK attempts=1, RetryTemplate retries=2 -> 3 request(s) for one call
  SDK attempts=3, RetryTemplate retries=0 -> 3 request(s) for one call
  SDK attempts=3, RetryTemplate retries=2 -> 9 request(s) for one call
  nothing configured at all              -> 5 requests for one call, and it took more than 5 seconds: true
With the retries set explicitly the counts are exactly attempts times (1 + retries). With nothing configured the call made five requests and took more than five seconds of retry waiting, all before the failover could start; on my machine that default took between about 25 and 70 seconds across runs, so I only assert the lower bound. The module sets both layers deliberately, putting the retry count on the Google client and zero on the Spring template, so a retry count means the same thing for all three providers.
Decide where retries live, once. Retry a little inside each provider (one or two, for blips) and let the fallback handle outages, or retry nothing and let the fallback do everything. Leaving the library defaults in place gives you a worst case you did not choose, and with Gemini that was a call lasting tens of seconds before any failover. Treat the request count in the transcript as the thing to check after every upgrade.
Going deeper: what failover does not give you
Failover moves a request; it does not make two providers interchangeable. Four things to plan for, none of which I tested. Answers differ: the same prompt will produce a different style, a different length and occasionally a different structure, so a feature that parses the reply (see structured output) should be run against every provider in the chain before you trust it. The fallback is cold: its prompt cache is empty, so the first requests after a failover pay the full input price, and a long outage on the primary can change your bill. Options do not travel: the options object you built for one provider is the wrong class for the next, as the options section showed, so a fallback chain needs its options set on each model, not per call. Streaming and tools add their own failure points mid-response, which a try-the-next-one wrapper around call does not cover. Section D of 05-fallback.txt shows what the caller sees when everything is down: the last provider’s error, with the whole trail kept for the log.

Should you even do this?

Keep the option open cheaply; run two providers only when you have a reason. Writing your code against ChatModel and building the model in one place costs almost nothing and lets you switch when pricing or quality changes. Running two providers in production, with failover, costs more: two sets of keys and quotas, prompts and parsers tested against both, and a decision about retries. Do it when an outage of one vendor would genuinely hurt, and measure before you commit to any claim about which is better for your task. What I did not test: any real vendor API, response quality, latency, streaming, tool calling across providers, the Vertex AI and cloud-marketplace routes to the same models, Gemini’s explicit cached-content feature, Anthropic’s one-hour cache and conversation-history strategies, and the property-based switch for all three. The requests, the usage plumbing and the retry counts are real; the tokens, the cache hits and the prices-to-dollars arithmetic are a worked example.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.