Add llm-gateway module: routing, failover below the tool-calling advisor, per-provider circuit breakers, dollar caps and token limits, tenant-keyed cache; real OpenAI and Anthropic models against a local fake
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,66 @@
|
||||
# llm-gateway
|
||||
|
||||
Companion code for [The LLM Gateway Pattern for Java Microservices](https://ankurm.com/llm-gateway-pattern-java-microservices/) and [LLM Gateway for Java Microservices: Complete Runnable Code](https://ankurm.com/llm-gateway-complete-example/), part of the [Spring AI series](../README.md) on ankurm.com.
|
||||
|
||||
A small Spring Boot 4.1 service that sits between your microservices and the LLM providers. A calling service sends one request and a tenant header. The gateway picks a provider chain from a model hint, fails over on transient errors, keeps one circuit breaker per provider, caps each tenant in tokens per hour and in dollars, and runs an allow-listed tool loop.
|
||||
|
||||
**No vendor API was called.** The tests use two kinds of stand-in, and each output file says which:
|
||||
|
||||
- `Scripted`, a `ChatModel` that plays back a script (answer, tool call, or an HTTP-style failure). Used where the point is the gateway's own logic.
|
||||
- `FakeVendors`, a local HTTP server that answers in the OpenAI and Anthropic wire formats. The **real** `OpenAiChatModel` and `AnthropicChatModel` talk to it, so the requests it records are the requests those classes would send, and the exceptions are the vendor SDKs' own.
|
||||
|
||||
Token counts in the fakes are numbers the test chose, and the semantic-cache test uses a hashed bag-of-words as its embedding model. Nothing here measures a real provider's latency, accuracy or cache behaviour.
|
||||
|
||||
## Versions
|
||||
|
||||
| Component | Version |
|
||||
|---|---|
|
||||
| Spring Boot | 4.1.1 (parent) |
|
||||
| Spring AI | 2.0.1 (`spring-ai-client-chat`, `spring-ai-openai`, `spring-ai-anthropic`) |
|
||||
| Resilience4j | 2.4.0 (`resilience4j-circuitbreaker`, used programmatically) |
|
||||
| Bucket4j | 8.21.0 (`bucket4j_jdk17-core`, in memory) |
|
||||
| Java | 25 (LTS) |
|
||||
|
||||
## Quickstart
|
||||
|
||||
```bash
|
||||
scripts/run-all.sh # runs the suite, regenerates output/
|
||||
OPENAI_API_KEY=... ANTHROPIC_API_KEY=... mvn spring-boot:run # a live gateway on :8080
|
||||
curl -X POST localhost:8080/v1/gateway/complete -H 'X-Tenant-Id: acme' -H 'Content-Type: application/json' \
|
||||
-d '{"featureTag":"support","userMessage":"What is your return policy?","modelHint":"smart","maxTokens":256}'
|
||||
```
|
||||
|
||||
Two consecutive test runs produce byte-identical files. The live run needs real keys and was not done for this article.
|
||||
|
||||
## What's here
|
||||
|
||||
| File | What it is |
|
||||
|---|---|
|
||||
| [`Failover.java`](src/main/java/com/ankurm/gateway/Failover.java) | The gateway's `ChatModel`: ordered targets, a breaker and a budget reservation per attempt |
|
||||
| [`Gateway.java`](src/main/java/com/ankurm/gateway/Gateway.java) | Rate limit, cache, `ChatClient` with `ToolCallingAdvisor`, metering |
|
||||
| [`Routes.java`](src/main/java/com/ankurm/gateway/Routes.java) | Hint to provider chain; an unknown hint is an error |
|
||||
| [`Budgets.java`](src/main/java/com/ankurm/gateway/Budgets.java) | Dollar caps: reserve before the call, settle after |
|
||||
| [`TokenLimiter.java`](src/main/java/com/ankurm/gateway/TokenLimiter.java) | Tokens per hour per tenant, Bucket4j |
|
||||
| [`SemanticCache.java`](src/main/java/com/ankurm/gateway/SemanticCache.java) | Embedding cache keyed by tenant and feature |
|
||||
| [`Failures.java`](src/main/java/com/ankurm/gateway/Failures.java) | HTTP status out of the vendor SDK exceptions |
|
||||
| [`GatewayConfig.java`](src/main/java/com/ankurm/gateway/GatewayConfig.java), [`application.yml`](src/main/resources/application.yml) | Providers, routes, tenants and breaker settings as data |
|
||||
|
||||
## Output files
|
||||
|
||||
| File | Written by |
|
||||
|---|---|
|
||||
| [`01-routing.txt`](output/01-routing.txt) | `RoutingTest`: hints, one model id per provider, the unknown-hint hazard |
|
||||
| [`02-failover.txt`](output/02-failover.txt) | `FailoverTest`: which statuses fail over |
|
||||
| [`03-circuit-breaker.txt`](output/03-circuit-breaker.txt) | `CircuitBreakerTest`: when a breaker opens, isolation, recovery |
|
||||
| [`04-tool-loop.txt`](output/04-tool-loop.txt) | `ToolLoopPlacementTest`: failover below vs above the tool advisor, usage totals |
|
||||
| [`05-wire.txt`](output/05-wire.txt) | `WireTest`: the real models failing over mid tool loop, the requests on the wire |
|
||||
| [`06-real-exceptions.txt`](output/06-real-exceptions.txt) | `WireTest`: what the SDKs throw for each status |
|
||||
| [`07-budget.txt`](output/07-budget.txt) | `BudgetTest`: caps, estimates, a cap hit mid-loop |
|
||||
| [`07b-budget-concurrency.txt`](output/07b-budget-concurrency.txt) | `BudgetTest`: 64 simultaneous requests |
|
||||
| [`08-rate-limit.txt`](output/08-rate-limit.txt) | `RateLimitTest`: pre-consume, settle, refill |
|
||||
| [`09-cache.txt`](output/09-cache.txt) | `CacheTest`: threshold and tenant key mechanics |
|
||||
| [`10-http.txt`](output/10-http.txt) | `HttpTest`: the application over HTTP |
|
||||
|
||||
## Not covered
|
||||
|
||||
Streaming (`ScopedValue` context does not cross Reactor threads), a shared store for breaker or limiter state across replicas (everything here is per instance), real provider behaviour, and anything measured about latency.
|
||||
Reference in New Issue
Block a user