Files
spring-ai/llm-gateway
..

llm-gateway

Companion code for The LLM Gateway Pattern for Java Microservices and LLM Gateway for Java Microservices: Complete Runnable Code, part of the Spring AI series on ankurm.com.

A small Spring Boot 4.1 service that sits between your microservices and the LLM providers. A calling service sends one request and a tenant header. The gateway picks a provider chain from a model hint, fails over on transient errors, keeps one circuit breaker per provider, caps each tenant in tokens per hour and in dollars, and runs an allow-listed tool loop.

No vendor API was called. The tests use two kinds of stand-in, and each output file says which:

  • Scripted, a ChatModel that plays back a script (answer, tool call, or an HTTP-style failure). Used where the point is the gateway's own logic.
  • FakeVendors, a local HTTP server that answers in the OpenAI and Anthropic wire formats. The real OpenAiChatModel and AnthropicChatModel talk to it, so the requests it records are the requests those classes would send, and the exceptions are the vendor SDKs' own.

Token counts in the fakes are numbers the test chose, and the semantic-cache test uses a hashed bag-of-words as its embedding model. Nothing here measures a real provider's latency, accuracy or cache behaviour.

Versions

Component Version
Spring Boot 4.1.1 (parent)
Spring AI 2.0.1 (spring-ai-client-chat, spring-ai-openai, spring-ai-anthropic)
Resilience4j 2.4.0 (resilience4j-circuitbreaker, used programmatically)
Bucket4j 8.21.0 (bucket4j_jdk17-core, in memory)
Java 25 (LTS)

Quickstart

scripts/run-all.sh                                   # runs the suite, regenerates output/
OPENAI_API_KEY=... ANTHROPIC_API_KEY=... mvn spring-boot:run   # a live gateway on :8080
curl -X POST localhost:8080/v1/gateway/complete -H 'X-Tenant-Id: acme' -H 'Content-Type: application/json' \
  -d '{"featureTag":"support","userMessage":"What is your return policy?","modelHint":"smart","maxTokens":256}'

Two consecutive test runs produce byte-identical files. The live run needs real keys and was not done for this article.

What's here

File What it is
Failover.java The gateway's ChatModel: ordered targets, a breaker and a budget reservation per attempt
Gateway.java Rate limit, cache, ChatClient with ToolCallingAdvisor, metering
Routes.java Hint to provider chain; an unknown hint is an error
Budgets.java Dollar caps: reserve before the call, settle after
TokenLimiter.java Tokens per hour per tenant, Bucket4j
SemanticCache.java Embedding cache keyed by tenant and feature
Failures.java HTTP status out of the vendor SDK exceptions
GatewayConfig.java, application.yml Providers, routes, tenants and breaker settings as data

Output files

File Written by
01-routing.txt RoutingTest: hints, one model id per provider, the unknown-hint hazard
02-failover.txt FailoverTest: which statuses fail over
03-circuit-breaker.txt CircuitBreakerTest: when a breaker opens, isolation, recovery
04-tool-loop.txt ToolLoopPlacementTest: failover below vs above the tool advisor, usage totals
05-wire.txt WireTest: the real models failing over mid tool loop, the requests on the wire
06-real-exceptions.txt WireTest: what the SDKs throw for each status
07-budget.txt BudgetTest: caps, estimates, a cap hit mid-loop
07b-budget-concurrency.txt BudgetTest: 64 simultaneous requests
08-rate-limit.txt RateLimitTest: pre-consume, settle, refill
09-cache.txt CacheTest: threshold and tenant key mechanics
10-http.txt HttpTest: the application over HTTP

Not covered

Streaming (ScopedValue context does not cross Reactor threads), a shared store for breaker or limiter state across replicas (everything here is per instance), real provider behaviour, and anything measured about latency.