Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
observability
Companion code for Observability for Spring AI: Tokens, Latency and Cost with Micrometer and OpenTelemetry, part of the Spring AI series on ankurm.com.
A small Spring AI app (/ask, /summarize, /weather with a tool call, /stream), the meters and spans the framework records for it with no code of ours, a 60-line handler that turns token usage into cost per endpoint, and a Grafana dashboard fed by Prometheus.
The model class is real, the provider is not. The app uses the real OpenAiChatModel and the real OpenAI Java SDK, pointed at FakeOpenAiServer, a local HTTP server that speaks the chat-completions protocol. So every meter name, tag, span and retry count is what Spring AI 2.0.1 really produces. But the token counts are about characters / 4, the latencies are the fake's, and the prices in application.yml are illustrative examples, not a current price list. No real provider was measured. The error test shows the SDK's retries against a server that always answers 500.
Versions
| Component | Version |
|---|---|
| Spring Boot | 4.1.1 |
| Spring AI | 2.0.1 |
| Micrometer / micrometer-tracing | 1.17.1 / 1.7.1 |
| OpenTelemetry | 1.62.0 |
| Prometheus / Grafana (release tarballs, no Docker) | 3.15.0 / 13.2.3 |
| Java | 25 (LTS) |
Quickstart
scripts/run-all.sh # 7 tests, then the stack, a load run and the dashboard check; regenerates output/01 .. 09
scripts/stack-up.sh # or just the stack: app :8080, Prometheus :9090, Grafana :3000 (admin/admin)
scripts/load.sh 20 # mixed traffic so every panel has data
scripts/stack-down.sh
Two consecutive runs of run-all.sh produce byte-identical output/ files. The Prometheus and Grafana tarballs must be unpacked under /tmp/tools/obs/{prom,grafana} (or set PROM and GRAFANA).
What's here
| File | What it is |
|---|---|
AiController.java, WeatherTools.java |
The endpoints and the one tool |
EndpointTagConvention.java |
Adds a low-cardinality app.endpoint tag to the framework's chat client meter |
UsageCostObservationHandler.java, Pricing.java |
Tokens and cost per endpoint, plus a counter for calls that could not be priced |
fake/FakeOpenAiServer.java |
The local stand-in for the provider (keywords slow, boom, nousage, weather) |
application.yml, application-stack.yml |
Settings; the stack profile adds latency histograms |
dashboards/spring-ai-observability.json |
The Grafana dashboard, generated by scripts/make-dashboard.py; screenshot.png is how it rendered |
stack/ |
Prometheus scrape config and Grafana provisioning |
Output files
| File | Written by |
|---|---|
01-builtin-metrics.txt |
BuiltInMetricsTest: the meters the framework creates |
02-tool-call-spans.txt |
ToolCallSpansTest: the span tree of a request with a tool call |
03-cost-per-endpoint.txt |
CostPerEndpointTest: tokens and cost per endpoint |
04-streaming-usage.txt |
StreamingUsageTest: usage on a streamed response |
05-error-and-retry.txt |
ErrorAndRetryTest: a failed call and the SDK's retries |
06-prompt-content.txt |
PromptContentTest: where prompt text appears, by default and when logging is on |
07-histogram.txt |
HistogramTest: latency buckets are opt-in |
08-dashboard-queries.txt |
scripts/verify-dashboard.py: every panel query through Prometheus and Grafana |
09-prometheus-families.txt |
The metric families as Prometheus sees them |