Files

4.4 KiB

observability

Companion code for Observability for Spring AI: Tokens, Latency and Cost with Micrometer and OpenTelemetry, part of the Spring AI series on ankurm.com.

A small Spring AI app (/ask, /summarize, /weather with a tool call, /stream), the meters and spans the framework records for it with no code of ours, a 60-line handler that turns token usage into cost per endpoint, and a Grafana dashboard fed by Prometheus.

The model class is real, the provider is not. The app uses the real OpenAiChatModel and the real OpenAI Java SDK, pointed at FakeOpenAiServer, a local HTTP server that speaks the chat-completions protocol. So every meter name, tag, span and retry count is what Spring AI 2.0.1 really produces. But the token counts are about characters / 4, the latencies are the fake's, and the prices in application.yml are illustrative examples, not a current price list. No real provider was measured. The error test shows the SDK's retries against a server that always answers 500.

Versions

Component Version
Spring Boot 4.1.1
Spring AI 2.0.1
Micrometer / micrometer-tracing 1.17.1 / 1.7.1
OpenTelemetry 1.62.0
Prometheus / Grafana (release tarballs, no Docker) 3.15.0 / 13.2.3
Java 25 (LTS)

Quickstart

scripts/run-all.sh      # 7 tests, then the stack, a load run and the dashboard check; regenerates output/01 .. 09
scripts/stack-up.sh     # or just the stack: app :8080, Prometheus :9090, Grafana :3000 (admin/admin)
scripts/load.sh 20      # mixed traffic so every panel has data
scripts/stack-down.sh

Two consecutive runs of run-all.sh produce byte-identical output/ files. The Prometheus and Grafana tarballs must be unpacked under /tmp/tools/obs/{prom,grafana} (or set PROM and GRAFANA).

What's here

File What it is
AiController.java, WeatherTools.java The endpoints and the one tool
EndpointTagConvention.java Adds a low-cardinality app.endpoint tag to the framework's chat client meter
UsageCostObservationHandler.java, Pricing.java Tokens and cost per endpoint, plus a counter for calls that could not be priced
fake/FakeOpenAiServer.java The local stand-in for the provider (keywords slow, boom, nousage, weather)
application.yml, application-stack.yml Settings; the stack profile adds latency histograms
dashboards/spring-ai-observability.json The Grafana dashboard, generated by scripts/make-dashboard.py; screenshot.png is how it rendered
stack/ Prometheus scrape config and Grafana provisioning

Output files

File Written by
01-builtin-metrics.txt BuiltInMetricsTest: the meters the framework creates
02-tool-call-spans.txt ToolCallSpansTest: the span tree of a request with a tool call
03-cost-per-endpoint.txt CostPerEndpointTest: tokens and cost per endpoint
04-streaming-usage.txt StreamingUsageTest: usage on a streamed response
05-error-and-retry.txt ErrorAndRetryTest: a failed call and the SDK's retries
06-prompt-content.txt PromptContentTest: where prompt text appears, by default and when logging is on
07-histogram.txt HistogramTest: latency buckets are opt-in
08-dashboard-queries.txt scripts/verify-dashboard.py: every panel query through Prometheus and Grafana
09-prometheus-families.txt The metric families as Prometheus sees them