Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
67 lines
4.8 KiB
Markdown
67 lines
4.8 KiB
Markdown
# llm-gateway
|
|
|
|
Companion code for [The LLM Gateway Pattern for Java Microservices](https://ankurm.com/llm-gateway-pattern-java-microservices/) and [LLM Gateway for Java Microservices: Complete Runnable Code](https://ankurm.com/llm-gateway-complete-example/), part of the [Spring AI series](../README.md) on ankurm.com.
|
|
|
|
A small Spring Boot 4.1 service that sits between your microservices and the LLM providers. A calling service sends one request and a tenant header. The gateway picks a provider chain from a model hint, fails over on transient errors, keeps one circuit breaker per provider, caps each tenant in tokens per hour and in dollars, and runs an allow-listed tool loop.
|
|
|
|
**No vendor API was called.** The tests use two kinds of stand-in, and each output file says which:
|
|
|
|
- `Scripted`, a `ChatModel` that plays back a script (answer, tool call, or an HTTP-style failure). Used where the point is the gateway's own logic.
|
|
- `FakeVendors`, a local HTTP server that answers in the OpenAI and Anthropic wire formats. The **real** `OpenAiChatModel` and `AnthropicChatModel` talk to it, so the requests it records are the requests those classes would send, and the exceptions are the vendor SDKs' own.
|
|
|
|
Token counts in the fakes are numbers the test chose, and the semantic-cache test uses a hashed bag-of-words as its embedding model. Nothing here measures a real provider's latency, accuracy or cache behaviour.
|
|
|
|
## Versions
|
|
|
|
| Component | Version |
|
|
|---|---|
|
|
| Spring Boot | 4.1.1 (parent) |
|
|
| Spring AI | 2.0.1 (`spring-ai-client-chat`, `spring-ai-openai`, `spring-ai-anthropic`) |
|
|
| Resilience4j | 2.4.0 (`resilience4j-circuitbreaker`, used programmatically) |
|
|
| Bucket4j | 8.21.0 (`bucket4j_jdk17-core`, in memory) |
|
|
| Java | 25 (LTS) |
|
|
|
|
## Quickstart
|
|
|
|
```bash
|
|
scripts/run-all.sh # runs the suite, regenerates output/
|
|
OPENAI_API_KEY=... ANTHROPIC_API_KEY=... mvn spring-boot:run # a live gateway on :8080
|
|
curl -X POST localhost:8080/v1/gateway/complete -H 'X-Tenant-Id: acme' -H 'Content-Type: application/json' \
|
|
-d '{"featureTag":"support","userMessage":"What is your return policy?","modelHint":"smart","maxTokens":256}'
|
|
```
|
|
|
|
Two consecutive test runs produce byte-identical files. The live run needs real keys and was not done for this article.
|
|
|
|
## What's here
|
|
|
|
| File | What it is |
|
|
|---|---|
|
|
| [`Failover.java`](src/main/java/com/ankurm/gateway/Failover.java) | The gateway's `ChatModel`: ordered targets, a breaker and a budget reservation per attempt |
|
|
| [`Gateway.java`](src/main/java/com/ankurm/gateway/Gateway.java) | Rate limit, cache, `ChatClient` with `ToolCallingAdvisor`, metering |
|
|
| [`Routes.java`](src/main/java/com/ankurm/gateway/Routes.java) | Hint to provider chain; an unknown hint is an error |
|
|
| [`Budgets.java`](src/main/java/com/ankurm/gateway/Budgets.java) | Dollar caps: reserve before the call, settle after |
|
|
| [`TokenLimiter.java`](src/main/java/com/ankurm/gateway/TokenLimiter.java) | Tokens per hour per tenant, Bucket4j |
|
|
| [`SemanticCache.java`](src/main/java/com/ankurm/gateway/SemanticCache.java) | Embedding cache keyed by tenant and feature |
|
|
| [`Failures.java`](src/main/java/com/ankurm/gateway/Failures.java) | HTTP status out of the vendor SDK exceptions |
|
|
| [`GatewayConfig.java`](src/main/java/com/ankurm/gateway/GatewayConfig.java), [`application.yml`](src/main/resources/application.yml) | Providers, routes, tenants and breaker settings as data |
|
|
|
|
## Output files
|
|
|
|
| File | Written by |
|
|
|---|---|
|
|
| [`01-routing.txt`](output/01-routing.txt) | `RoutingTest`: hints, one model id per provider, the unknown-hint hazard |
|
|
| [`02-failover.txt`](output/02-failover.txt) | `FailoverTest`: which statuses fail over |
|
|
| [`03-circuit-breaker.txt`](output/03-circuit-breaker.txt) | `CircuitBreakerTest`: when a breaker opens, isolation, recovery |
|
|
| [`04-tool-loop.txt`](output/04-tool-loop.txt) | `ToolLoopPlacementTest`: failover below vs above the tool advisor, usage totals |
|
|
| [`05-wire.txt`](output/05-wire.txt) | `WireTest`: the real models failing over mid tool loop, the requests on the wire |
|
|
| [`06-real-exceptions.txt`](output/06-real-exceptions.txt) | `WireTest`: what the SDKs throw for each status |
|
|
| [`07-budget.txt`](output/07-budget.txt) | `BudgetTest`: caps, estimates, a cap hit mid-loop |
|
|
| [`07b-budget-concurrency.txt`](output/07b-budget-concurrency.txt) | `BudgetTest`: 64 simultaneous requests |
|
|
| [`08-rate-limit.txt`](output/08-rate-limit.txt) | `RateLimitTest`: pre-consume, settle, refill |
|
|
| [`09-cache.txt`](output/09-cache.txt) | `CacheTest`: threshold and tenant key mechanics |
|
|
| [`10-http.txt`](output/10-http.txt) | `HttpTest`: the application over HTTP |
|
|
|
|
## Not covered
|
|
|
|
Streaming (`ScopedValue` context does not cross Reactor threads), a shared store for breaker or limiter state across replicas (everything here is per instance), real provider behaviour, and anything measured about latency.
|