Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
56 lines
4.4 KiB
Markdown
56 lines
4.4 KiB
Markdown
# observability
|
|
|
|
Companion code for [Observability for Spring AI: Tokens, Latency and Cost with Micrometer and OpenTelemetry](https://ankurm.com/spring-ai-2-0-observability-micrometer-opentelemetry-tokens-cost/), part of the [Spring AI series](../README.md) on ankurm.com.
|
|
|
|
A small Spring AI app (`/ask`, `/summarize`, `/weather` with a tool call, `/stream`), the meters and spans the framework records for it with no code of ours, a 60-line handler that turns token usage into cost per endpoint, and a Grafana dashboard fed by Prometheus.
|
|
|
|
**The model class is real, the provider is not.** The app uses the real `OpenAiChatModel` and the real OpenAI Java SDK, pointed at `FakeOpenAiServer`, a local HTTP server that speaks the chat-completions protocol. So every meter name, tag, span and retry count is what Spring AI 2.0.1 really produces. But the token counts are about characters / 4, the latencies are the fake's, and the prices in `application.yml` are illustrative examples, not a current price list. No real provider was measured. The error test shows the SDK's retries against a server that always answers 500.
|
|
|
|
## Versions
|
|
|
|
| Component | Version |
|
|
|---|---|
|
|
| Spring Boot | 4.1.1 |
|
|
| Spring AI | 2.0.1 |
|
|
| Micrometer / micrometer-tracing | 1.17.1 / 1.7.1 |
|
|
| OpenTelemetry | 1.62.0 |
|
|
| Prometheus / Grafana (release tarballs, no Docker) | 3.15.0 / 13.2.3 |
|
|
| Java | 25 (LTS) |
|
|
|
|
## Quickstart
|
|
|
|
```bash
|
|
scripts/run-all.sh # 7 tests, then the stack, a load run and the dashboard check; regenerates output/01 .. 09
|
|
scripts/stack-up.sh # or just the stack: app :8080, Prometheus :9090, Grafana :3000 (admin/admin)
|
|
scripts/load.sh 20 # mixed traffic so every panel has data
|
|
scripts/stack-down.sh
|
|
```
|
|
|
|
Two consecutive runs of `run-all.sh` produce byte-identical `output/` files. The Prometheus and Grafana tarballs must be unpacked under `/tmp/tools/obs/{prom,grafana}` (or set `PROM` and `GRAFANA`).
|
|
|
|
## What's here
|
|
|
|
| File | What it is |
|
|
|---|---|
|
|
| [`AiController.java`](src/main/java/com/ankurm/observability/AiController.java), [`WeatherTools.java`](src/main/java/com/ankurm/observability/WeatherTools.java) | The endpoints and the one tool |
|
|
| [`EndpointTagConvention.java`](src/main/java/com/ankurm/observability/EndpointTagConvention.java) | Adds a low-cardinality `app.endpoint` tag to the framework's chat client meter |
|
|
| [`UsageCostObservationHandler.java`](src/main/java/com/ankurm/observability/UsageCostObservationHandler.java), [`Pricing.java`](src/main/java/com/ankurm/observability/Pricing.java) | Tokens and cost per endpoint, plus a counter for calls that could not be priced |
|
|
| [`fake/FakeOpenAiServer.java`](src/main/java/com/ankurm/observability/fake/FakeOpenAiServer.java) | The local stand-in for the provider (keywords `slow`, `boom`, `nousage`, `weather`) |
|
|
| [`application.yml`](src/main/resources/application.yml), [`application-stack.yml`](src/main/resources/application-stack.yml) | Settings; the `stack` profile adds latency histograms |
|
|
| [`dashboards/spring-ai-observability.json`](dashboards/spring-ai-observability.json) | The Grafana dashboard, generated by [`scripts/make-dashboard.py`](scripts/make-dashboard.py); [`screenshot.png`](dashboards/screenshot.png) is how it rendered |
|
|
| [`stack/`](stack) | Prometheus scrape config and Grafana provisioning |
|
|
|
|
## Output files
|
|
|
|
| File | Written by |
|
|
|---|---|
|
|
| [`01-builtin-metrics.txt`](output/01-builtin-metrics.txt) | `BuiltInMetricsTest`: the meters the framework creates |
|
|
| [`02-tool-call-spans.txt`](output/02-tool-call-spans.txt) | `ToolCallSpansTest`: the span tree of a request with a tool call |
|
|
| [`03-cost-per-endpoint.txt`](output/03-cost-per-endpoint.txt) | `CostPerEndpointTest`: tokens and cost per endpoint |
|
|
| [`04-streaming-usage.txt`](output/04-streaming-usage.txt) | `StreamingUsageTest`: usage on a streamed response |
|
|
| [`05-error-and-retry.txt`](output/05-error-and-retry.txt) | `ErrorAndRetryTest`: a failed call and the SDK's retries |
|
|
| [`06-prompt-content.txt`](output/06-prompt-content.txt) | `PromptContentTest`: where prompt text appears, by default and when logging is on |
|
|
| [`07-histogram.txt`](output/07-histogram.txt) | `HistogramTest`: latency buckets are opt-in |
|
|
| [`08-dashboard-queries.txt`](output/08-dashboard-queries.txt) | `scripts/verify-dashboard.py`: every panel query through Prometheus and Grafana |
|
|
| [`09-prometheus-families.txt`](output/09-prometheus-families.txt) | The metric families as Prometheus sees them |
|