Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
llm-gateway
Companion code for The LLM Gateway Pattern for Java Microservices and LLM Gateway for Java Microservices: Complete Runnable Code, part of the Spring AI series on ankurm.com.
A small Spring Boot 4.1 service that sits between your microservices and the LLM providers. A calling service sends one request and a tenant header. The gateway picks a provider chain from a model hint, fails over on transient errors, keeps one circuit breaker per provider, caps each tenant in tokens per hour and in dollars, and runs an allow-listed tool loop.
No vendor API was called. The tests use two kinds of stand-in, and each output file says which:
Scripted, aChatModelthat plays back a script (answer, tool call, or an HTTP-style failure). Used where the point is the gateway's own logic.FakeVendors, a local HTTP server that answers in the OpenAI and Anthropic wire formats. The realOpenAiChatModelandAnthropicChatModeltalk to it, so the requests it records are the requests those classes would send, and the exceptions are the vendor SDKs' own.
Token counts in the fakes are numbers the test chose, and the semantic-cache test uses a hashed bag-of-words as its embedding model. Nothing here measures a real provider's latency, accuracy or cache behaviour.
Versions
| Component | Version |
|---|---|
| Spring Boot | 4.1.1 (parent) |
| Spring AI | 2.0.1 (spring-ai-client-chat, spring-ai-openai, spring-ai-anthropic) |
| Resilience4j | 2.4.0 (resilience4j-circuitbreaker, used programmatically) |
| Bucket4j | 8.21.0 (bucket4j_jdk17-core, in memory) |
| Java | 25 (LTS) |
Quickstart
scripts/run-all.sh # runs the suite, regenerates output/
OPENAI_API_KEY=... ANTHROPIC_API_KEY=... mvn spring-boot:run # a live gateway on :8080
curl -X POST localhost:8080/v1/gateway/complete -H 'X-Tenant-Id: acme' -H 'Content-Type: application/json' \
-d '{"featureTag":"support","userMessage":"What is your return policy?","modelHint":"smart","maxTokens":256}'
Two consecutive test runs produce byte-identical files. The live run needs real keys and was not done for this article.
What's here
| File | What it is |
|---|---|
Failover.java |
The gateway's ChatModel: ordered targets, a breaker and a budget reservation per attempt |
Gateway.java |
Rate limit, cache, ChatClient with ToolCallingAdvisor, metering |
Routes.java |
Hint to provider chain; an unknown hint is an error |
Budgets.java |
Dollar caps: reserve before the call, settle after |
TokenLimiter.java |
Tokens per hour per tenant, Bucket4j |
SemanticCache.java |
Embedding cache keyed by tenant and feature |
Failures.java |
HTTP status out of the vendor SDK exceptions |
GatewayConfig.java, application.yml |
Providers, routes, tenants and breaker settings as data |
Output files
| File | Written by |
|---|---|
01-routing.txt |
RoutingTest: hints, one model id per provider, the unknown-hint hazard |
02-failover.txt |
FailoverTest: which statuses fail over |
03-circuit-breaker.txt |
CircuitBreakerTest: when a breaker opens, isolation, recovery |
04-tool-loop.txt |
ToolLoopPlacementTest: failover below vs above the tool advisor, usage totals |
05-wire.txt |
WireTest: the real models failing over mid tool loop, the requests on the wire |
06-real-exceptions.txt |
WireTest: what the SDKs throw for each status |
07-budget.txt |
BudgetTest: caps, estimates, a cap hit mid-loop |
07b-budget-concurrency.txt |
BudgetTest: 64 simultaneous requests |
08-rate-limit.txt |
RateLimitTest: pre-consume, settle, refill |
09-cache.txt |
CacheTest: threshold and tenant key mechanics |
10-http.txt |
HttpTest: the application over HTTP |
Not covered
Streaming (ScopedValue context does not cross Reactor threads), a shared store for breaker or limiter state across replicas (everything here is per instance), real provider behaviour, and anything measured about latency.