Spring Boot startup time: bean-by-bean diagnosis, and one directory per post

Adds spring-boot-startup-time/, the companion project for BLOG-618: a runnable
Spring Boot 4.1.1 application on JDK 25 that installs BufferingApplicationStartup
and FlightRecorderApplicationStartup behind a system property, and a /diag/startup
endpoint that computes step self time -- the number /actuator/startup does not give
you and the one that names the actual culprits.

Captured under docs/output/: the step tree sorted both ways, the same startup as JFR
events, a +5000-class experiment putting 0.11 ms per scanned class on the classpath
scan tax, the silent truncation a 2048-step buffer performs, and JDK 25 AOT cache
timings (6.93 s to 4.82 s). Post body and metadata live in post/.

Moves the existing Actuator project into actuator-in-production/ so the repository
holds one directory per article; the root README is now an index.
This commit is contained in:
2026-09-04 23:58:38 +05:30
parent 4b6cefa60a
commit 958b401f0f
112 changed files with 2744 additions and 154 deletions

169
actuator-in-production/README.md Executable file
View File

@@ -0,0 +1,169 @@
# spring-boot-demo — Actuator in production
Companion repository for **[Spring Boot Actuator in Production](https://ankurm.com/spring-boot-actuator-production-endpoints-security-health-indicators/)** on ankurm.com.
Every status code, JSON body, byte count and timing figure in the article was produced by
running this project. The transcripts live in [`docs/output/`](docs/output) and are regenerated
by a single command.
---
## Versions
Verified with `mvn dependency:list` on the machine that produced `docs/output/` — see
[`docs/output/00-versions.txt`](docs/output/00-versions.txt).
| Component | Version | Notes |
|---|---|---|
| Spring Boot | 4.1.1 | GA 20 August 2026 |
| Spring Framework | 7.0.9 | via `spring-boot-starter-parent` |
| Spring Security | 7.1.1 | via `spring-boot-starter-parent` |
| `spring-boot-actuator` | 4.1.1 | |
| `spring-boot-health` | 4.1.1 | **new module in Boot 4**`HealthIndicator` lives here now |
| `spring-boot-restclient` | 4.1.1 | **not pulled in by the web starter** |
| Micrometer | 1.17.1 | there is no Micrometer 2.x GA; see [`docs/02`](docs/02-boot-4-changes.md) |
| `kafka-clients` | 4.2.1 | what Boot 4.1.1 manages (4.3.1 is the latest on Central) |
| H2 | 2.4.240 | |
| JDK | Temurin 25.0.4.1+1 LTS | |
| Maven | 3.9.11 | |
---
## Quickstart
```bash
export JAVA_HOME=/path/to/jdk-25
mvn -DskipTests package
# defaults: only /actuator/health is exposed
java -jar target/actuator-production-1.0.0.jar
# every endpoint, no authentication - the configuration you should never ship
java -jar target/actuator-production-1.0.0.jar --spring.profiles.active=exposeall,open
# the configuration you should ship
java -jar target/actuator-production-1.0.0.jar --spring.profiles.active=secured
```
Credentials for every profile that requires them: **`ops` / `ops-password`**.
Regenerate every transcript in `docs/output/`:
```bash
JAVA_HOME=/path/to/jdk-25 ./scripts/run-all.sh
```
That takes roughly three minutes, most of which is the naive Kafka scenario blocking for its
full 60 seconds. That is the point of it.
---
## Profiles
| Profile | What it demonstrates |
|---|---|
| *(none)* | Boot defaults. Only `health` on the web, `show-details: never` |
| `exposeall` | `management.endpoints.web.exposure.include: "*"` |
| `open` | A `permitAll` security chain — the misconfiguration, kept on purpose |
| `secured` | `EndpointRequest.toAnyEndpoint()` + `ROLE_ACTUATOR` + `when-authorized` details |
| `mgmtport` | Actuator on port 9001, base path `/manage`, bound to loopback |
| `details` | `show-details: always` — the full component breakdown |
| `groups` | `liveness` / `readiness` / `startup` groups wired correctly |
| `kafkanaive` | The textbook `AdminClient` health check, so its 60-second block can be timed |
Profiles compose: `--spring.profiles.active=exposeall,open`.
---
## Endpoints
Read from the running application, not from the documentation — see
[`docs/output/02-endpoint-catalogue.txt`](docs/output/02-endpoint-catalogue.txt).
| Endpoint | Web-exposed by default | `access` default | Notes |
|---|---|---|---|
| `health` | **yes** | `unrestricted` | the only one exposed out of the box |
| `info` | no | `unrestricted` | Boot 4.1 added `process.*` fields |
| `beans` | no | `unrestricted` | 426 beans in this app, with types and wiring |
| `conditions` | no | `unrestricted` | the auto-configuration report |
| `configprops` | no | `unrestricted` | values masked like `env` |
| `env` | no | `unrestricted` | **masks every value** unless `show-values` says otherwise |
| `loggers` | no | `unrestricted` | has a `POST` — a write endpoint |
| `mappings` | no | `unrestricted` | every URL your app serves |
| `metrics` | no | `unrestricted` | from `spring-boot-micrometer-metrics` |
| `prometheus` | no | `unrestricted` | needs `micrometer-registry-prometheus` |
| `sbom` | no | `unrestricted` | |
| `scheduledtasks` | no | `unrestricted` | |
| `threaddump` | no | `unrestricted` | two operations: JSON and `text/plain` |
| `heapdump` | no | **`none`** | `include: "*"` is **not** enough |
| `shutdown` | no | **`none`** | `include: "*"` is **not** enough |
| `startup` | no | `unrestricted` | needs a `BufferingApplicationStartup` |
| `httpexchanges` | no | `unrestricted` | needs an `HttpExchangeRepository` bean |
| `auditevents` | no | `unrestricted` | needs an `AuditEventRepository` bean |
| `logfile` | no | `unrestricted` | needs `logging.file.name` |
| `caches`, `flyway`, `liquibase`, `quartz`, `sessions`, `integrationgraph` | no | `unrestricted` | conditional on the relevant module |
| `diag` | no | `unrestricted` | **this repository's own** — delete before shipping |
`heapdump` and `shutdown` are the only two endpoints whose `access` defaults to `none`. That
list came from Spring Boot's own `spring-configuration-metadata.json`, not from a blog.
---
## Documentation
| # | Chapter |
|---|---|
| 01 | [What Actuator actually exposes](docs/01-what-actuator-exposes.md) |
| 02 | [What changed in Spring Boot 4](docs/02-boot-4-changes.md) |
| 03 | [The endpoint catalogue](docs/03-endpoint-catalogue.md) |
| 04 | [Securing Actuator](docs/04-securing-actuator.md) |
| 05 | [Custom health indicators](docs/05-custom-health-indicators.md) |
| 06 | [Health indicator failure modes](docs/06-health-indicator-failure-modes.md) |
| 07 | [Groups, probes and Kubernetes](docs/07-groups-and-probes.md) |
| 08 | [The diagnostics endpoint](docs/08-diagnostics.md) |
| 09 | [Testing Actuator](docs/09-testing-actuator.md) |
---
## Captured output
| File | Scenario |
|---|---|
| [`00-versions.txt`](docs/output/00-versions.txt) | resolved dependency versions |
| [`01-default-exposure.txt`](docs/output/01-default-exposure.txt) | Actuator with zero configuration |
| [`02-endpoint-catalogue.txt`](docs/output/02-endpoint-catalogue.txt) | every exposed endpoint, from the running app |
| [`03-open-actuator-leak.txt`](docs/output/03-open-actuator-leak.txt) | what an anonymous caller really gets |
| [`04-heapdump-leak.txt`](docs/output/04-heapdump-leak.txt) | 59 MB, plaintext credentials inside |
| [`05-secured-matrix.txt`](docs/output/05-secured-matrix.txt) | the full authorisation matrix |
| [`06-management-port.txt`](docs/output/06-management-port.txt) | port 9001, and what it isolates |
| [`07-custom-health-indicators.txt`](docs/output/07-custom-health-indicators.txt) | DB, Kafka and external API |
| [`08-groups-and-probes.txt`](docs/output/08-groups-and-probes.txt) | liveness 200 while readiness 503 |
| [`09-kafka-timeout.txt.tuned`](docs/output/09-kafka-timeout.txt.tuned) | 1.6 s |
| [`09-kafka-timeout.txt.naive`](docs/output/09-kafka-timeout.txt.naive) | **60.2 s** |
| [`10-slow-upstream.txt`](docs/output/10-slow-upstream.txt) | a read timeout doing its job |
Status codes and bodies are reproducible. Timing figures are indicative and drift between
machines — except the 60-second one, which is a Kafka default and lands on 60.0 s every time.
---
## Layout
```
pom.xml
scripts/
run-all.sh regenerate everything below docs/output/
run.sh / stop.sh start and stop with given profiles
demo-*.sh one script per captured scenario
src/main/java/com/ankurm/actuator/
health/ OrdersDatabase, Kafka, ExternalApi indicators + UpstreamState
config/ Secured, Open and baseline security chains
web/ DiagnosticsEndpoint, stub upstream, business controller
src/test/java/ contract tests for the surprising behaviour
docs/ numbered chapters
docs/output/ captured real output
```
## Licence
MIT — see [LICENSE](../LICENSE).

View File

@@ -0,0 +1,57 @@
[← README](../README.md) · **01 · What Actuator actually exposes** · [02 What changed in Spring Boot 4 →](02-boot-4-changes.md)
# 01 — What Actuator actually exposes
Add the starter, configure nothing, and ask the application what it publishes.
```xml
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-actuator</artifactId>
</dependency>
```
The answer, from [`../docs/output/01-default-exposure.txt`](output/01-default-exposure.txt):
```json
{"_links":{"self":{...},"health":{...},"health-path":{...}}}
```
One endpoint. `GET /actuator/env` returns **404**, not 403 — it was discovered, it exists as a
JMX endpoint, and it is simply not mapped onto HTTP.
That distinction matters more than it looks:
- **Discovery** — Actuator finds every `@Endpoint` bean on the classpath.
- **Access** (`management.endpoint.<id>.access`) — whether the endpoint may be operated at all.
Defaults to `unrestricted` for everything except `heapdump` and `shutdown`.
- **Exposure** (`management.endpoints.web.exposure.include`) — whether it is mapped onto HTTP.
Defaults to `health` only.
All three have to line up. A 404 from an Actuator path tells you nothing about whether the
endpoint is enabled, and people read it as "it's off" when it is often "it's on, over JMX".
## The default health body
```json
{"groups":["liveness","readiness"],"status":"DOWN"}
```
Two things to notice.
**The groups are there by default.** In Spring Boot 3 the liveness and readiness probes only
appeared when you asked for them or when Boot detected Kubernetes. In Boot 4 they are enabled
out of the box — see [chapter 02](02-boot-4-changes.md).
**`show-details` defaults to `never`**, so even an authenticated caller sees a bare status. That
is a sensible default and almost everyone overrides it to `always` without thinking about who
can reach the endpoint. [Chapter 04](04-securing-actuator.md) covers the middle option.
**The status is `DOWN`** because this repository registers a Kafka indicator and there is no
broker. One custom indicator that touches a third party is all it takes to turn the default
`/actuator/health` red — and that URL is what most Kubernetes manifests point their readiness
*and* liveness probes at. [Chapter 07](07-groups-and-probes.md) is about not doing that.
---
[← README](../README.md) · **01** · [02 What changed in Spring Boot 4 →](02-boot-4-changes.md)

View File

@@ -0,0 +1,107 @@
[← 01 What Actuator exposes](01-what-actuator-exposes.md) · **02 · What changed in Spring Boot 4** · [03 The endpoint catalogue →](03-endpoint-catalogue.md)
# 02 — What changed in Spring Boot 4
Every claim below was checked against the 4.1.1 jars with `javap`, or against the 3.5.16 jars
for the "before" side. None of it came from a migration blog.
## `HealthIndicator` moved module and package
This is the one that breaks every custom health indicator ever written.
```
Boot 3.5.16 org.springframework.boot.actuate.health.HealthIndicator (spring-boot-actuator)
Boot 4.1.1 org.springframework.boot.health.contributor.HealthIndicator (spring-boot-health)
```
`org.springframework.boot.actuate.health` does not exist in Boot 4.1.1 at all — not deprecated,
absent. `Health`, `Status`, `AbstractHealthIndicator`, `HealthContributor` and the composites
all moved with it, into a new `spring-boot-health` module.
Two neighbouring packages in the same new module are worth knowing:
- `org.springframework.boot.health.application``DiskSpaceHealthIndicator`,
`LivenessStateHealthIndicator`, `ReadinessStateHealthIndicator`, `SslHealthIndicator`
- `org.springframework.boot.health.registry``HealthContributorRegistry`, which is how
[`DiagnosticsEndpoint`](../src/main/java/com/ankurm/actuator/web/DiagnosticsEndpoint.java)
enumerates contributors at runtime
## The interface method was renamed
```java
// Boot 3.5.16
public interface HealthIndicator extends HealthContributor {
default Health getHealth(boolean includeDetails);
Health health();
}
// Boot 4.1.1
public interface HealthIndicator extends HealthContributor {
default Health health(boolean includeDetails);
Health health();
}
```
If you overrode `getHealth(boolean)` — which people do to control detail rendering — it now
silently stops being an override. Add `@Override` and let the compiler find it.
`Health` also stopped extending `HealthComponent`; that class is gone from Boot 4.1.1 entirely.
## `EndpointRequest` moved too
```
Boot 3.5.16 org.springframework.boot.actuate.autoconfigure.security.servlet.EndpointRequest
Boot 4.1.1 org.springframework.boot.security.autoconfigure.actuate.web.servlet.EndpointRequest
```
Same class, same static methods (`toAnyEndpoint()`, `to(Class...)`, `to(String...)`,
`toLinks()`, `toAdditionalPaths(...)`), new package, and it now ships in `spring-boot-security`
rather than `spring-boot-actuator-autoconfigure`. The reactive variant made the equivalent move.
## `RestClient` is not in the web starter any more
Boot 4 split the framework into fine-grained modules, and `RestClient` auto-configuration went
with it. Adding `spring-boot-starter-webmvc` and injecting `RestClient.Builder` fails at
startup:
```
Parameter 0 of constructor in com.ankurm.actuator.health.ExternalApiHealthIndicator
required a bean of type 'org.springframework.web.client.RestClient$Builder' that could not be found.
```
The fix is one dependency:
```xml
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-restclient</artifactId>
</dependency>
```
This cost a startup cycle while building this repository, which is why it is documented here
rather than glossed over.
## Liveness and readiness probes are on by default
From the 4.0 migration guide, and confirmed by running the app with no configuration at all —
`/actuator/health` reports `"groups":["liveness","readiness"]` out of the box. Turn them off
with `management.endpoint.health.probes.enabled=false` if you genuinely do not want them.
## Nullability annotations
Actuator endpoint parameters can no longer use `org.springframework.lang.Nullable` to mark a
parameter optional. Migrate to `org.jspecify.annotations.Nullable`.
## There is no Micrometer 2
Several migration write-ups claim Boot 4 replaces "legacy Actuator endpoints" with a
"Micrometer 2 observability stack". Spring Boot 4.1.1 resolves **Micrometer 1.17.1**
(`docs/output/00-versions.txt`), and `maven-metadata.xml` for `micrometer-core` shows 1.17.1 as
the newest GA with 1.18.0-M1 as a milestone. Metrics *did* move to their own module
(`spring-boot-micrometer-metrics`, which `spring-boot-starter-actuator` still pulls in), and
there is a new `spring-boot-starter-opentelemetry`. Neither of those is a major Micrometer
release.
---
[← 01](01-what-actuator-exposes.md) · **02** · [03 The endpoint catalogue →](03-endpoint-catalogue.md)

View File

@@ -0,0 +1,89 @@
[← 02 Boot 4 changes](02-boot-4-changes.md) · **03 · The endpoint catalogue** · [04 Securing Actuator →](04-securing-actuator.md)
# 03 — The endpoint catalogue
The table in the [README](../README.md#endpoints) is the reference. This chapter covers the
three endpoints people get wrong.
## `env` does not leak values
The oldest Actuator scare story is "`/actuator/env` dumps your database password". In Spring
Boot 4 it does not. From [`output/03-open-actuator-leak.txt`](output/03-open-actuator-leak.txt),
with no credentials sent at all:
```json
{"source":"Config resource 'class path resource [application.yaml]' ...","value":"******"}
```
That is `spring.datasource.password`, and it is also `acme.partner.credential` — a property
whose name matches none of the classic `password`/`secret`/`token`/`key` patterns. Both masked.
Masking is not pattern matching on key names. It is `management.endpoint.env.show-values`,
which defaults to `never`. Set it to `always` and both come back in plaintext:
```
acme.partner.credential S3CRET-partner-credential
spring.datasource.password not-a-real-password-but-watch-what-/actuator/env-does-with-it
spring.security.user.password ops-password
```
So the risk with `env` is real but it is a *configuration* risk, not a default. The `origin`
field is also worth knowing about — it reports the exact file and line a property came from,
which is genuinely the fastest way to answer "where is this value coming from" in a service
with six property sources.
## `heapdump` is not exposed by `include: "*"`
This surprises people who learned Actuator on Boot 2.
```
$ curl -s -o /dev/null -w '%{http_code}' http://localhost:8080/actuator/heapdump
404
```
…with `management.endpoints.web.exposure.include: "*"` and a `permitAll` security chain.
`management.endpoint.heapdump.access` defaults to `none`, and exposure does not override
access. Reading Spring Boot's own `spring-configuration-metadata.json` for every
`management.endpoint.*.access` key gives exactly two endpoints defaulting to `none`:
| Endpoint | `access` default |
|---|---|
| `heapdump` | `none` |
| `shutdown` | `none` |
| everything else | `unrestricted` |
Turn it on and the story changes completely —
[`output/04-heapdump-leak.txt`](output/04-heapdump-leak.txt):
```
status=200 bytes=59065395 type=application/octet-stream
$ strings /tmp/heap.hprof | grep -c 'S3CRET-partner-credential'
1
$ strings /tmp/heap.hprof | grep -o 'not-a-real-password[^"]*' | head -1
not-a-real-password-but-watch-what-/actuator/env-does-with-it
```
59 MB containing, in plaintext, both properties that `/actuator/env` had just masked.
Sanitisation is a rendering feature of one endpoint. It is not a security boundary.
## `loggers` is a write endpoint
`GET /actuator/loggers/{name}` is harmless. `POST` is not:
```
$ curl -X POST -H 'Content-Type: application/json' \
-d '{"configuredLevel":"TRACE"}' /actuator/loggers/org.springframework
status=204
```
204, no credentials. An attacker who can do that has a denial-of-service primitive (TRACE on a
busy service will fill a disk) and, depending on your logging configuration, a way to get
request bodies and headers written to a file they may be able to read through another channel.
`beans` (426 entries here) and `mappings` (30 servlet mappings) are read-only but are pure
reconnaissance — they describe your entire application's shape.
---
[← 02](02-boot-4-changes.md) · **03** · [04 Securing Actuator →](04-securing-actuator.md)

View File

@@ -0,0 +1,122 @@
[← 03 Endpoint catalogue](03-endpoint-catalogue.md) · **04 · Securing Actuator** · [05 Custom health indicators →](05-custom-health-indicators.md)
# 04 — Securing Actuator
Two mechanisms, and you want both. Network isolation decides who can reach the port;
authorisation decides who can use it. Neither is a substitute for the other.
## The separate management port
[`SecuredActuatorConfig`](../src/main/java/com/ankurm/actuator/config/SecuredActuatorConfig.java)
handles authorisation. `application-mgmtport.yaml` handles reachability:
```yaml
management:
server:
port: 9001
address: 127.0.0.1
endpoints:
web:
base-path: /manage
```
From [`output/06-management-port.txt`](output/06-management-port.txt):
```
GET :8080/actuator/health 401 <- application chain, no Actuator here
GET :8080/orders/count 200
GET :9001/manage/health 503
GET :9001/actuator/health 404 <- base-path moved
GET :9001/orders/count 404 <- separate context, no app controllers
```
```
LISTEN *:8080
LISTEN [::ffff:127.0.0.1]:9001
```
The last two lines are the argument. The management port runs a **separate application context**
with its own `DispatcherServlet`, so it cannot serve your controllers, and binding it to
loopback (or simply not listing it as a Kubernetes Service port) makes Actuator unreachable
from the internet by routing rather than by a rule somebody has to keep correct.
This is defence in depth, not a replacement for authorisation. Anything running in the same pod
or on the same host still reaches 9001.
## Match on `EndpointRequest`, not on a path
```java
http.securityMatcher(EndpointRequest.toAnyEndpoint())
.authorizeHttpRequests((requests) -> requests
.requestMatchers(EndpointRequest.to(HealthEndpoint.class, InfoEndpoint.class)).permitAll()
.anyRequest().hasRole("ACTUATOR"))
```
A rule written as `requestMatchers("/actuator/**")` stops matching the moment somebody sets
`management.endpoints.web.base-path` — as the `mgmtport` profile does. The endpoint moves, the
rule does not follow, and the endpoint is now unprotected with no error anywhere. `EndpointRequest`
asks the endpoint registry where things actually are.
## Order the chains
`@Order(1)` on the Actuator chain, `@Order(2)` on the application chain. Without explicit
ordering, whichever chain Spring registers first wins for a given request, and an application
chain ending in `permitAll()` will happily swallow `/actuator/**`.
## The resulting matrix
From [`output/05-secured-matrix.txt`](output/05-secured-matrix.txt), with exposure set to `"*"`:
| Request | Anonymous | `ops` / ROLE_ACTUATOR | Wrong password |
|---|---|---|---|
| `/actuator/health` | 503 | 503 | — |
| `/actuator/info` | 200 | — | — |
| `/actuator/env` | 401 | 200 | 401 |
| `/actuator/beans` | 401 | 200 | — |
| `/actuator/threaddump` | 401 | 200 | — |
| `/actuator` (links) | 401 | — | — |
| `/orders/count` | 200 | — | — |
Exposure is `"*"` and nothing leaks. That is the point: **exposure is not access control.**
Use exposure to decide what exists, and the security chain to decide who may use it.
## `show-details: when-authorized`
The third option, and the one to reach for. Anonymous:
```json
{"groups":["liveness","readiness"],"status":"DOWN"}
```
Authenticated with `ROLE_ACTUATOR`: the full nine-component breakdown, naming `kafka` as the
failing one. Same URL, same status code, different body. Your load balancer gets the signal it
needs; an anonymous prober learns that something is wrong but not which of your dependencies to
attack next.
```yaml
management:
endpoint:
health:
show-details: when-authorized
roles: ACTUATOR
```
## Make the Actuator chain stateless and CSRF-free
Spring Boot's auto-configured security is browser-shaped: it enables CSRF protection. Actuator
is a machine-to-machine API, and `POST /actuator/loggers/{name}` from a script will be rejected
before it reaches the endpoint.
While building this repository that showed up as a `401` on `POST /stub/upstream/mode` while
`GET /orders/count` with the *same credentials* returned `200` — a genuinely confusing pair of
results that has nothing to do with the credentials. See
[`DemoSecurityConfig`](../src/main/java/com/ankurm/actuator/config/DemoSecurityConfig.java).
```java
.csrf((csrf) -> csrf.disable())
.sessionManagement((s) -> s.sessionCreationPolicy(SessionCreationPolicy.STATELESS))
```
---
[← 03](03-endpoint-catalogue.md) · **04** · [05 Custom health indicators →](05-custom-health-indicators.md)

View File

@@ -0,0 +1,112 @@
[← 04 Securing Actuator](04-securing-actuator.md) · **05 · Custom health indicators** · [06 Failure modes →](06-health-indicator-failure-modes.md)
# 05 — Custom health indicators
Three indicators, three different jobs.
| Bean | Class | Checks |
|---|---|---|
| `ordersDatabase` | [`OrdersDatabaseHealthIndicator`](../src/main/java/com/ankurm/actuator/health/OrdersDatabaseHealthIndicator.java) | the query the application depends on |
| `kafka` | [`KafkaHealthIndicator`](../src/main/java/com/ankurm/actuator/health/KafkaHealthIndicator.java) | `AdminClient.describeCluster` |
| `externalApi` | [`ExternalApiHealthIndicator`](../src/main/java/com/ankurm/actuator/health/ExternalApiHealthIndicator.java) | an HTTP dependency, with a budget |
## The shape
```java
import org.springframework.boot.health.contributor.AbstractHealthIndicator;
import org.springframework.boot.health.contributor.Health;
@Component("ordersDatabase")
public class OrdersDatabaseHealthIndicator extends AbstractHealthIndicator {
@Override
protected void doHealthCheck(Health.Builder builder) throws Exception {
// throw, or call builder.up() / down() / status("DEGRADED")
}
}
```
Note the package — `org.springframework.boot.health.contributor`, new in Boot 4. See
[chapter 02](02-boot-4-changes.md).
`AbstractHealthIndicator` over the bare interface, because it catches your exceptions and turns
them into `DOWN` with the error recorded, rather than letting one throwing indicator take down
the whole `/actuator/health` response.
## The bean name is the JSON key
`@Component("ordersDatabase")` produces `"ordersDatabase"` in the response. Boot strips a
trailing `HealthIndicator` from the bean name, so `ordersDatabaseHealthIndicator` gives the same
key. Rename the bean and every dashboard, alert and health-group `include:` that referenced the
old key silently stops matching — which is why
[`HealthGroupTests`](../src/test/java/com/ankurm/actuator/HealthGroupTests.java) asserts on the
names.
## Check what the application needs, not what the connection can do
Boot's built-in `db` indicator runs a validation query and answers "can I reach the database".
That is rarely the question. `ordersDatabase` runs the query the service actually depends on and
adds a latency judgement:
```json
{"details":{"orders":3,"queryMs":0,"slowThresholdMs":250},"status":"UP"}
```
It returns a custom `DEGRADED` status above the threshold. Custom statuses are legal;
`management.endpoint.health.status.order` and `.http-mapping` control how they aggregate and
what HTTP code they produce. An unmapped custom status aggregates as `UNKNOWN` and maps to 200,
which is usually not what you meant.
## Every indicator needs a budget
All three carry an explicit timeout, and the reason is [chapter 06](06-health-indicator-failure-modes.md).
- `ordersDatabase``jdbc.setQueryTimeout(2)`
- `externalApi` — connect and read timeouts of 750 ms on the `RestClient` request factory
- `kafka` — four separate Kafka bounds plus an outer `KafkaFuture.get(timeout)`
## Credentials
`externalApi` sends HTTP Basic. The first captured run of this repository did not, and reported:
```json
{"error":"HttpClientErrorException$Unauthorized: 401 : [no body]","status":"DOWN"}
```
which looks exactly like an upstream outage and was a missing header. If your indicator
authenticates, make the failure message distinguish "they are down" from "we are unauthorised" —
future you will be reading it at 3am.
## Reuse the client
`KafkaHealthIndicator` creates one `AdminClient` in its constructor and closes it in
`close()`. Creating one per probe opens a fresh set of broker connections on every probe
interval; across a fleet of any size that is a denial-of-service against your own brokers.
## What it all looks like
[`output/07-custom-health-indicators.txt`](output/07-custom-health-indicators.txt) has the full
nine-component response. Flipping the stub upstream to `DOWN` and changing nothing else:
```json
{
"status": "DOWN",
"externalApi": {
"details": {
"error": "HttpServerErrorException$ServiceUnavailable: 503 : \"upstream unavailable\"",
"url": "http://localhost:8080/stub/upstream/ping",
"latencyMs": 101,
"timeoutMs": 750
},
"status": "DOWN"
}
}
```
`/actuator/health` now answers 503. If that URL is your readiness probe, every pod in the
deployment just left the load balancer because a third party had a bad minute.
[Chapter 07](07-groups-and-probes.md) fixes that.
---
[← 04](04-securing-actuator.md) · **05** · [06 Failure modes →](06-health-indicator-failure-modes.md)

View File

@@ -0,0 +1,100 @@
[← 05 Custom health indicators](05-custom-health-indicators.md) · **06 · Health indicator failure modes** · [07 Groups and probes →](07-groups-and-probes.md)
# 06 — Health indicator failure modes
## The 60-second Kafka health check
This is the most useful number in the repository. The textbook Kafka health check:
```java
AdminClient admin = AdminClient.create(Map.of(BOOTSTRAP_SERVERS_CONFIG, "localhost:9092"));
admin.describeCluster().nodes().get();
```
Against an unreachable broker, from
[`output/09-kafka-timeout.txt.naive`](output/09-kafka-timeout.txt.naive):
```json
{"details":{"tuned":false,"error":"TimeoutException: Timed out waiting for a node assignment. Call: listNodes","probeMs":60002,"budgetMs":180000},"status":"DOWN"}
wall clock: 60.2 s
```
**60.002 seconds.** The outer budget was 180 s, so it was not the binding constraint — Kafka's
own `default.api.timeout.ms` was. These are the `AdminClientConfig` defaults, read from
`AdminClientConfig.configDef()` on kafka-clients 4.2.1:
| Property | Default |
|---|---|
| `request.timeout.ms` | 30000 |
| `default.api.timeout.ms` | **60000** |
| `socket.connection.setup.timeout.ms` | 10000 |
| `retries` | 2147483647 |
| `metadata.recovery.strategy` | `rebootstrap` |
The trap is that `request.timeout.ms` is the one everybody sets. It bounds a single attempt, and
with `retries` at `Integer.MAX_VALUE` the client simply attempts again until
`default.api.timeout.ms` fires. Set only `request.timeout.ms` and your health check still blocks
for a minute.
The tuned version, same broker, same absence of it —
[`output/09-kafka-timeout.txt.tuned`](output/09-kafka-timeout.txt.tuned):
```json
{"details":{"tuned":true,"error":"TimeoutException: null","probeMs":1500,"budgetMs":1500},"status":"DOWN"}
wall clock: 1.6 s
```
Forty times faster, from setting four properties and one `KafkaFuture.get(timeout)`.
Why this matters beyond the number: your probe interval is probably 10 seconds. A check that
takes 60 seconds means six probes are in flight at once, each holding a container worker thread,
for as long as the broker is down. That is how a Kafka outage becomes an application outage.
## The rebootstrap storm
Kafka 4 defaults `metadata.recovery.strategy` to `rebootstrap`. A long-lived `AdminClient`
pointed at a dead broker keeps a background thread retrying, and it logs. The first run of this
repository produced **192** `Rebootstrapping with Cluster(id = null, ...)` lines in a few
seconds, and the sandbox ran out of memory shortly after.
`metadata.recovery.strategy: none` plus a sane `reconnect.backoff.max.ms` stops it. Whatever you
choose, know that the AdminClient's background thread is doing work whether or not anyone is
calling your health endpoint.
## The unbounded HTTP call
`ExternalApiHealthIndicator` sets connect and read timeouts of 750 ms. Against a stub that
sleeps 30 seconds — [`output/10-slow-upstream.txt`](output/10-slow-upstream.txt):
```json
{"details":{"error":"ResourceAccessException: I/O error on GET request ...: Read timed out","latencyMs":753,"timeoutMs":750},"status":"DOWN"}
wall clock: 0.84 s
```
`latencyMs: 753` against `timeoutMs: 750`. Remove `setReadTimeout` and that becomes 30 seconds,
holding a Tomcat worker the entire time.
The default for `SimpleClientHttpRequestFactory` — and for `HttpURLConnection` underneath it —
is no timeout at all. Not a long timeout. None.
## Rules of thumb
- **Every remote call in a health indicator needs an explicit timeout**, and the total of all of
them should be comfortably under your probe interval.
- **Set the API-level timeout, not just the request-level one.** Kafka is the sharpest example
but the pattern recurs — clients with internal retry loops need an outer bound.
- **Fail fast and fail loudly.** A `DOWN` with a useful `error` detail beats a probe that hangs.
- **Health checks should be cheap.** They run on every instance on every probe interval, forever.
- **Do not check dependencies you cannot act on.** If a third party being down does not change
what this instance should do, do not put it in a health indicator at all.
There is a cache if you need it: `management.endpoint.health.cache.time-to-live`. Reach for it
only after you have fixed the timeouts — caching a 60-second check gives you a fast endpoint
serving a stale answer, which is worse than a slow honest one.
---
[← 05](05-custom-health-indicators.md) · **06** · [07 Groups and probes →](07-groups-and-probes.md)

View File

@@ -0,0 +1,130 @@
[← 06 Failure modes](06-health-indicator-failure-modes.md) · **07 · Groups, probes and Kubernetes** · [08 The diagnostics endpoint →](08-diagnostics.md)
# 07 — Groups, probes and Kubernetes
## The failure this prevents
A third-party API goes down. Your `externalApi` indicator reports `DOWN`. `/actuator/health`
aggregates to `DOWN` and answers 503. Your Kubernetes manifest points **both** probes at it:
```yaml
livenessProbe: { httpGet: { path: /actuator/health, port: 8080 } }
readinessProbe: { httpGet: { path: /actuator/health, port: 8080 } }
```
The readiness failure is correct — this instance cannot serve. The liveness failure is a
catastrophe: the kubelet kills the container, every replica fails the same probe at the same
time, and the whole deployment enters a restart loop. A partial outage in someone else's system
has become a total outage in yours, and the restarts make recovery slower because every fresh
JVM has to warm up while the upstream is still failing.
**Liveness answers "is this process broken beyond recovery?"** Almost nothing external belongs
in it, because restarting cannot fix an external problem.
**Readiness answers "should this instance receive traffic right now?"** Dependencies belong here.
## The configuration
From `application-groups.yaml`:
```yaml
management:
endpoint:
health:
group:
liveness:
include: livenessState,diskSpace
readiness:
include: readinessState,ordersDatabase,externalApi
startup:
include: ordersDatabase
probes:
enabled: true
```
Each group gets its own path: `/actuator/health/liveness`, `/actuator/health/readiness`,
`/actuator/health/startup`. Groups can carry their own `show-details` too.
## The evidence
[`output/08-groups-and-probes.txt`](output/08-groups-and-probes.txt):
```
--- baseline: upstream UP ---
GET /actuator/health HTTP 503
GET /actuator/health/liveness HTTP 200
GET /actuator/health/readiness HTTP 200
GET /actuator/health/startup HTTP 200
--- upstream DOWN ---
GET /actuator/health HTTP 503
GET /actuator/health/liveness HTTP 200
GET /actuator/health/readiness HTTP 503
GET /actuator/health/startup HTTP 200
```
Liveness stayed 200 across the outage. Readiness moved. Kubernetes removes this instance from
the Service endpoints and leaves the process alone; when the upstream recovers, readiness goes
green and traffic returns without a single restart.
Note the aggregate `/actuator/health` was 503 in *both* columns — the Kafka indicator is down
throughout, since this repository runs no broker. That is exactly why you should not point a
probe at the aggregate: it is the union of everything, and it tells the orchestrator nothing
actionable.
## The manifest
```yaml
livenessProbe:
httpGet: { path: /actuator/health/liveness, port: 9001 }
periodSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet: { path: /actuator/health/readiness, port: 9001 }
periodSeconds: 5
failureThreshold: 2
startupProbe:
httpGet: { path: /actuator/health/startup, port: 9001 }
periodSeconds: 5
failureThreshold: 30
```
Port 9001 is the management port from [chapter 04](04-securing-actuator.md). The kubelet reaches
it inside the pod network; the ingress does not.
The `startupProbe` matters more than people think. Without one, a JVM that takes 45 seconds to
warm up under `failureThreshold: 3` and `periodSeconds: 10` gets killed at 30 seconds, forever,
and the symptom is a crash-loop with no error in the logs.
## `readinessState`, and shutting down cleanly
`readinessState` is Boot's own indicator, driven by `ApplicationAvailability`. During graceful
shutdown Boot flips it to `REFUSING_TRAFFIC` before the server stops accepting, so readiness
goes red while in-flight requests finish. Pair it with:
```yaml
server:
shutdown: graceful
spring:
lifecycle:
timeout-per-shutdown-phase: 30s
```
You can also drive it yourself by publishing an `AvailabilityChangeEvent` — useful for taking an
instance out of rotation before a risky migration.
## Which dependencies belong in readiness
Ask: *if this is down, can this instance still serve any useful request?*
- Primary database → yes, readiness.
- Kafka, for a service whose only job is consuming → yes, readiness.
- Kafka, for a service that also serves reads → probably not; degrade rather than withdraw.
- A recommendations API you fall back to a static list for → no. Do not check it at all.
Every dependency in readiness is a dependency that can take your service out of rotation. That
list should be shorter than your instinct suggests.
---
[← 06](06-health-indicator-failure-modes.md) · **07** · [08 The diagnostics endpoint →](08-diagnostics.md)

View File

@@ -0,0 +1,95 @@
[← 07 Groups and probes](07-groups-and-probes.md) · **08 · The diagnostics endpoint** · [09 Testing Actuator →](09-testing-actuator.md)
# 08 — The diagnostics endpoint
[`DiagnosticsEndpoint`](../src/main/java/com/ankurm/actuator/web/DiagnosticsEndpoint.java) is a
custom `@Endpoint(id = "diag")` that prints the Actuator state no configuration file will tell
you: which endpoints were actually discovered and mapped onto HTTP, what paths and methods each
publishes, and which health contributors are registered.
**Delete it before you ship.** The list of exposed endpoints is itself reconnaissance.
## Why it exists
Reading `application.yaml` and predicting exposure is how people ship `/actuator/heapdump` to
the internet. `management.endpoints.web.exposure.include: "*"` looks like it exposes everything;
it does not expose `heapdump` or `shutdown` ([chapter 03](03-endpoint-catalogue.md)). Conditional
endpoints like `httpexchanges` and `startup` are absent unless a specific bean exists. Ask the
running application instead of predicting.
## Writing a custom endpoint
```java
@Component
@Endpoint(id = "diag")
public class DiagnosticsEndpoint {
@ReadOperation
public Map<String, Object> diagnostics() { ... }
}
```
`@ReadOperation` maps to GET, `@WriteOperation` to POST, `@DeleteOperation` to DELETE. The id
must be lowercase alphanumeric; it becomes the path segment and the JMX name. You still have to
expose it — `include: health,info,diag` or `"*"`.
The interesting part is the three beans it injects:
| Bean | Gives you |
|---|---|
| `WebEndpointsSupplier` | every `ExposableWebEndpoint`, with its `WebOperation` predicates |
| `HealthContributorRegistry` | the live contributor tree, composites included |
| `Environment` | the effective values of the `management.*` keys |
`HealthContributorRegistry` is in `org.springframework.boot.health.registry` — new in Boot 4
([chapter 02](02-boot-4-changes.md)). It iterates as `HealthContributors.Entry`, and nested
`HealthContributors` are composites, which is why the collection is recursive.
## The output
From [`output/02-endpoint-catalogue.txt`](output/02-endpoint-catalogue.txt):
```json
{
"activeProfiles": ["exposeall", "open"],
"serverPort": "8080",
"managementPort": "(same as server.port)",
"managementBasePath": "/actuator",
"exposureInclude": "*",
"healthShowDetails": "always",
"exposedWebEndpointCount": 14,
"exposedWebEndpoints": {
"loggers": ["GET loggers", "GET loggers/{name}", "POST loggers/{name}"],
"threaddump": ["GET threaddump", "GET threaddump"],
...
},
"healthContributors": [
"db (DataSourceHealthIndicator)",
"diskSpace (DiskSpaceHealthIndicator)",
"externalApi (ExternalApiHealthIndicator)",
"kafka (KafkaHealthIndicator)",
"livenessState (LivenessStateHealthIndicator)",
"ordersDatabase (OrdersDatabaseHealthIndicator)",
"ping (PingHealthIndicator)",
"readinessState (ReadinessStateHealthIndicator)",
"ssl (SslHealthIndicator)"
]
}
```
Two things worth pausing on.
`"loggers"` shows the `POST` operation explicitly. That is the write endpoint from
[chapter 03](03-endpoint-catalogue.md), visible in a list rather than remembered.
`"threaddump"` appears **twice** with the same path. It publishes two operations that differ
only in what they produce — `application/json` and `text/plain`. Content negotiation, not a bug.
Run the same endpoint under the `mgmtport` profile and `managementPort` reads `9001`,
`managementBasePath` reads `/manage`, and the endpoint count is unchanged — the endpoints moved,
they did not disappear. That is the fact that breaks path-string security rules
([chapter 04](04-securing-actuator.md)).
---
[← 07](07-groups-and-probes.md) · **08** · [09 Testing Actuator →](09-testing-actuator.md)

View File

@@ -0,0 +1,95 @@
[← 08 The diagnostics endpoint](08-diagnostics.md) · **09 · Testing Actuator** · [README →](../README.md)
# 09 — Testing Actuator
Eight tests, and none of them assert the happy path. They pin the behaviour that would be
expensive to rediscover after an upgrade.
## Pin the surprises
[`ActuatorExposureTests`](../src/test/java/com/ankurm/actuator/ActuatorExposureTests.java) runs
with `exposeall,secured` — the widest exposure with real security — and asserts:
```java
@Test
void heapdumpIsNotExposedEvenWithWildcardExposure() throws Exception {
mvc().perform(get("/actuator/heapdump")).andExpect(status().isNotFound());
}
```
If a future Boot release changes the `access` default for `heapdump`, or if somebody adds
`management.endpoint.heapdump.access=unrestricted` to a shared config, this test is where you
find out. Same for `shutdown`.
```java
@Test
void envMasksEveryValueRegardlessOfKeyName() throws Exception {
// acme.partner.credential matches no password/secret/token pattern.
... .andExpect(jsonPath("$.property.value").value("******"));
}
@Test
void anonymousHealthCarriesNoComponentBreakdown() throws Exception {
mvc().perform(get("/actuator/health"))
.andExpect(jsonPath("$.components").doesNotExist());
}
```
That last one is the `show-details: when-authorized` contract from
[chapter 04](04-securing-actuator.md), asserted rather than assumed. It is a one-line
configuration change away from leaking your dependency map to anonymous callers.
## Test the operational contract, not the indicator
[`HealthGroupTests`](../src/test/java/com/ankurm/actuator/HealthGroupTests.java):
```java
@Test
void upstreamOutageMovesTheExternalApiIndicatorButNotLiveness() {
upstream.set(UpstreamState.Mode.DOWN);
assertThat(statusOf("externalApi")).isEqualTo(Status.DOWN);
assertThat(statusOf("livenessState")).isEqualTo(Status.UP);
}
```
That is [chapter 07](07-groups-and-probes.md) as an executable assertion. It fails if someone
adds `externalApi` to the liveness group, which is the exact change that causes a fleet-wide
restart loop.
```java
@Test
void theExternalApiIndicatorRespectsItsTimeoutBudget() {
upstream.set(UpstreamState.Mode.SLOW); // the stub sleeps 30s
...
assertThat(elapsedMs).isLessThan(5_000L);
}
```
Fails if someone removes the read timeout. A generous bound, because timing assertions on CI are
how you get flaky suites — but 5 s versus 30 s is a wide enough gap to be safe.
## `DEFINED_PORT`, and why
```java
@SpringBootTest(webEnvironment = SpringBootTest.WebEnvironment.DEFINED_PORT)
```
Not the default `MOCK` environment. `ExternalApiHealthIndicator` makes a real HTTP call to this
application's own stub controller, so a servlet container has to be listening on the port it was
configured with. Under `MOCK` the indicator reports `DOWN` with a connection error and every
assertion about upstream state is meaningless — which is precisely what happened on the first
run of this suite.
`RANDOM_PORT` does not help here: the indicator resolves its URL at bean construction, before
the server binds, so `local.server.port` is not available to it.
## What is not tested
There is no test for the 60-second Kafka timeout. A test that takes a minute to pass does not
belong in a suite people run before pushing —
[`scripts/demo-kafka-timeout.sh`](../scripts/demo-kafka-timeout.sh) captures it instead, and the
transcript is committed. Some evidence is better as a script than as an assertion.
---
[← 08](08-diagnostics.md) · **09** · [README →](../README.md)

View File

@@ -0,0 +1,23 @@
$ java -version
openjdk version "25.0.4.1" 2026-08-18 LTS
OpenJDK Runtime Environment Temurin-25.0.4.1+1 (build 25.0.4.1+1-LTS)
OpenJDK 64-Bit Server VM Temurin-25.0.4.1+1 (build 25.0.4.1+1-LTS, mixed mode, sharing)
$ mvn -v | head -3
Apache Maven 3.9.11 (3e54c93a704957b63ee3494413a2b544fd3d825b)
Maven home: /tmp/tools/apache-maven-3.9.11
Java version: 25.0.4.1, vendor: Eclipse Adoptium, runtime: /tmp/tools/jdk-25.0.4.1+1
$ mvn dependency:list -- the resolved versions behind spring-boot-starter-parent 4.1.1
com.h2database:h2 2.4.240
io.micrometer:micrometer-core 1.17.1
io.micrometer:micrometer-registry-prometheus 1.17.1
org.apache.kafka:kafka-clients 4.2.1
org.springframework.boot:spring-boot 4.1.1
org.springframework.boot:spring-boot-actuator 4.1.1
org.springframework.boot:spring-boot-health 4.1.1
org.springframework.boot:spring-boot-restclient 4.1.1
org.springframework.security:spring-security-core 7.1.1
org.springframework.security:spring-security-web 7.1.1
org.springframework:spring-core 7.0.9
org.springframework:spring-web 7.0.9

View File

@@ -0,0 +1,35 @@
### Spring Boot Actuator, starter added, ZERO management.* configuration
$ curl -s -u ops:ops-password http://localhost:8080/actuator
{
"_links": {
"self": {
"href": "http://localhost:8080/actuator",
"templated": false
},
"health": {
"href": "http://localhost:8080/actuator/health",
"templated": false
},
"health-path": {
"href": "http://localhost:8080/actuator/health/{*path}",
"templated": true
}
}
}
$ curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password http://localhost:8080/actuator/env
404
404 = discovered but NOT exposed over HTTP. Exposure and existence are different things.
$ curl -s -u ops:ops-password http://localhost:8080/actuator/health
{
"groups": [
"liveness",
"readiness"
],
"status": "DOWN"
}
Only 'health' is web-exposed by default. show-details defaults to 'never', so even an
authenticated caller sees a bare status until you say otherwise.

View File

@@ -0,0 +1,80 @@
### Every web-exposed endpoint, read from the running application
### profiles: exposeall,open management.endpoints.web.exposure.include: "*"
$ curl -s http://localhost:8080/actuator/diag
{
"activeProfiles": [
"exposeall",
"open"
],
"serverPort": "8080",
"managementPort": "(same as server.port)",
"managementBasePath": "/actuator",
"exposureInclude": "*",
"exposureExclude": "(none)",
"healthShowDetails": "always",
"exposedWebEndpointCount": 14,
"exposedWebEndpoints": {
"beans": [
"GET beans"
],
"conditions": [
"GET conditions"
],
"configprops": [
"GET configprops",
"GET configprops/{prefix}"
],
"diag": [
"GET diag"
],
"env": [
"GET env",
"GET env/{toMatch}"
],
"health": [
"GET health",
"GET health/{*path}"
],
"info": [
"GET info"
],
"loggers": [
"GET loggers",
"GET loggers/{name}",
"POST loggers/{name}"
],
"mappings": [
"GET mappings"
],
"metrics": [
"GET metrics",
"GET metrics/{requiredMetricName}"
],
"prometheus": [
"GET prometheus"
],
"sbom": [
"GET sbom",
"GET sbom/{id}"
],
"scheduledtasks": [
"GET scheduledtasks"
],
"threaddump": [
"GET threaddump",
"GET threaddump"
]
},
"healthContributors": [
"db (DataSourceHealthIndicator)",
"diskSpace (DiskSpaceHealthIndicator)",
"externalApi (ExternalApiHealthIndicator)",
"kafka (KafkaHealthIndicator)",
"livenessState (LivenessStateHealthIndicator)",
"ordersDatabase (OrdersDatabaseHealthIndicator)",
"ping (PingHealthIndicator)",
"readinessState (ReadinessStateHealthIndicator)",
"ssl (SslHealthIndicator)"
]
}

View File

@@ -0,0 +1,39 @@
### profiles: exposeall,open -- NO credentials are sent on any request below
--- 1. /actuator/env does NOT leak values in Spring Boot 4 ---
$ curl -s http://localhost:8080/actuator/env/spring.datasource.password | jq .property
{
"source": "Config resource 'class path resource [application.yaml]' via location 'optional:classpath:/'",
"value": "******"
}
$ curl -s http://localhost:8080/actuator/env/acme.partner.credential | jq .property
{
"source": "Config resource 'class path resource [application.yaml]' via location 'optional:classpath:/'",
"value": "******"
}
Note the second one. 'acme.partner.credential' matches none of the classic
password/secret/token key patterns, and it is still masked. Masking is driven by
management.endpoint.env.show-values, which defaults to 'never' - not by key names.
--- 2. /actuator/heapdump is NOT exposed by 'include: "*"' ---
$ curl -s -o /dev/null -w '%{http_code}' http://localhost:8080/actuator/heapdump
404
management.endpoint.heapdump.access defaults to 'none'. So does shutdown.
They are the only two endpoints that do.
--- 3. /actuator/loggers: an unauthenticated WRITE ---
$ curl -s -X POST -d '{"configuredLevel":"TRACE"}' -H 'Content-Type: application/json' http://localhost:8080/actuator/loggers/org.springframework
status=204
$ curl -s http://localhost:8080/actuator/loggers/org.springframework
{
"configuredLevel": "TRACE",
"effectiveLevel": "TRACE"
}
(level reset). An attacker who can flip your root logger to TRACE has both a
denial-of-service primitive and a way to get request bodies written to disk.
--- 4. /actuator/beans and /actuator/mappings: your whole application, described ---
$ curl -s http://localhost:8080/actuator/mappings | python3 -c 'count the URL patterns'
30 servlet mappings disclosed
426 beans disclosed, each with its type and dependencies

View File

@@ -0,0 +1,13 @@
### profiles: exposeall,open PLUS --management.endpoint.heapdump.access=unrestricted
$ curl -s -o /tmp/heap.hprof -w 'status=%{http_code} bytes=%{size_download} type=%{content_type}' http://localhost:8080/actuator/heapdump
status=200 bytes=59186852 type=application/octet-stream
$ strings /tmp/heap.hprof | grep -c 'S3CRET-partner-credential'
1
$ strings /tmp/heap.hprof | grep -o 'not-a-real-password[^"]*' | head -1
not-a-real-password-but-watch-what-/actuator/env-does-with-it!
/actuator/env masked both of these to ******.
/actuator/heapdump handed over the process memory that contains them in plaintext.
Sanitisation is a property-rendering feature. It is not a security boundary.

View File

@@ -0,0 +1,106 @@
### profile: secured (SecuredActuatorConfig + application-secured.yaml)
### exposure is "*" - the security chain, not the exposure list, is what protects it
ANONYMOUS
GET /actuator/health 503
GET /actuator/info 200
GET /actuator/env 401
GET /actuator/beans 401
GET /actuator/threaddump 401
GET /actuator (the links index) 401
AUTHENTICATED as ops (ROLE_ACTUATOR)
GET /actuator/health 503
GET /actuator/env 200
GET /actuator/beans 200
GET /actuator/threaddump 200
WRONG PASSWORD
GET /actuator/env 401
--- health body, anonymous (show-details: when-authorized) ---
{
"groups": [
"liveness",
"readiness"
],
"status": "DOWN"
}
--- health body, authenticated as ROLE_ACTUATOR ---
{
"components": {
"db": {
"details": {
"database": "H2",
"validationQuery": "isValid()"
},
"status": "UP"
},
"diskSpace": {
"details": {
"total": 10213466112,
"free": 3877920768,
"threshold": 10485760,
"path": "/tmp/work/spring-boot-demo/.",
"exists": true
},
"status": "UP"
},
"externalApi": {
"details": {
"url": "http://localhost:8080/stub/upstream/ping",
"response": "pong",
"latencyMs": 4,
"timeoutMs": 750
},
"status": "UP"
},
"kafka": {
"details": {
"bootstrap": "localhost:9092",
"tuned": true,
"error": "TimeoutException: Timed out waiting for a node assignment. Call: listNodes",
"probeMs": 1500,
"budgetMs": 1500
},
"status": "DOWN"
},
"livenessState": {
"status": "UP"
},
"ordersDatabase": {
"details": {
"orders": 3,
"queryMs": 0,
"slowThresholdMs": 250
},
"status": "UP"
},
"ping": {
"status": "UP"
},
"readinessState": {
"status": "UP"
},
"ssl": {
"details": {
"expiringChains": [],
"invalidChains": [],
"validChains": []
},
"status": "UP"
}
},
"groups": [
"liveness",
"readiness"
],
"status": "DOWN"
}
Same endpoint, same status code, different body. An anonymous prober learns that the
service is unhealthy but not WHICH dependency is unhealthy.
--- the business endpoint is untouched by the actuator chain ---
GET /orders/count (anonymous) 200

View File

@@ -0,0 +1,101 @@
### profile: mgmtport
### management.server.port: 9001 / management.server.address: 127.0.0.1 / base-path: /manage
--- the application port no longer serves Actuator at all ---
GET :8080/actuator/health 401
GET :8080/manage/health 401
GET :8080/orders/count 200
--- the management port serves it on the new base path ---
GET :9001/manage/health 503
GET :9001/actuator/health 404
GET :9001/orders/count 404
Note the last line. The management context has its own DispatcherServlet and does NOT
see application controllers. That is the isolation you are paying for.
--- what the management context reports about itself ---
$ curl -s -u ops:ops-password http://127.0.0.1:9001/manage/diag
{
"activeProfiles": [
"mgmtport"
],
"serverPort": "8080",
"managementPort": "9001",
"managementBasePath": "/manage",
"exposureInclude": "*",
"exposureExclude": "(none)",
"healthShowDetails": "never",
"exposedWebEndpointCount": 14,
"exposedWebEndpoints": {
"beans": [
"GET beans"
],
"conditions": [
"GET conditions"
],
"configprops": [
"GET configprops",
"GET configprops/{prefix}"
],
"diag": [
"GET diag"
],
"env": [
"GET env",
"GET env/{toMatch}"
],
"health": [
"GET health",
"GET health/{*path}"
],
"info": [
"GET info"
],
"loggers": [
"GET loggers",
"GET loggers/{name}",
"POST loggers/{name}"
],
"mappings": [
"GET mappings"
],
"metrics": [
"GET metrics",
"GET metrics/{requiredMetricName}"
],
"prometheus": [
"GET prometheus"
],
"sbom": [
"GET sbom",
"GET sbom/{id}"
],
"scheduledtasks": [
"GET scheduledtasks"
],
"threaddump": [
"GET threaddump",
"GET threaddump"
]
},
"healthContributors": [
"db (DataSourceHealthIndicator)",
"diskSpace (DiskSpaceHealthIndicator)",
"externalApi (ExternalApiHealthIndicator)",
"kafka (KafkaHealthIndicator)",
"livenessState (LivenessStateHealthIndicator)",
"ordersDatabase (OrdersDatabaseHealthIndicator)",
"ping (PingHealthIndicator)",
"readinessState (ReadinessStateHealthIndicator)",
"ssl (SslHealthIndicator)"
]
}
--- listening sockets ---
$ ss -ltn | grep -E ':(8080|9001)'
LISTEN 0 100 *:8080 *:*
LISTEN 0 100 [::ffff:127.0.0.1]:9001 *:*
9001 is bound to 127.0.0.1 only. 8080 is bound to *. An ingress that forwards to 8080
cannot reach Actuator no matter how the security rules are written.

View File

@@ -0,0 +1,121 @@
### profile: details (show-details: always, show-components: always)
--- upstream UP ---
$ curl -s -o /dev/null -w '%{http_code}' http://localhost:8080/actuator/health
HTTP 503
{
"components": {
"db": {
"details": {
"database": "H2",
"validationQuery": "isValid()"
},
"status": "UP"
},
"diskSpace": {
"details": {
"total": 10213466112,
"free": 3877675008,
"threshold": 10485760,
"path": "/tmp/work/spring-boot-demo/.",
"exists": true
},
"status": "UP"
},
"externalApi": {
"details": {
"url": "http://localhost:8080/stub/upstream/ping",
"response": "pong",
"latencyMs": 66,
"timeoutMs": 750
},
"status": "UP"
},
"kafka": {
"details": {
"bootstrap": "localhost:9092",
"tuned": true,
"error": "TimeoutException: null",
"probeMs": 1500,
"budgetMs": 1500
},
"status": "DOWN"
},
"livenessState": {
"status": "UP"
},
"ordersDatabase": {
"details": {
"orders": 3,
"queryMs": 0,
"slowThresholdMs": 250
},
"status": "UP"
},
"ping": {
"status": "UP"
},
"readinessState": {
"status": "UP"
},
"ssl": {
"details": {
"expiringChains": [],
"invalidChains": [],
"validChains": []
},
"status": "UP"
}
},
"groups": [
"liveness",
"readiness"
],
"status": "DOWN"
}
--- flip the upstream to DOWN, change nothing else ---
$ curl -s -X POST 'http://localhost:8080/stub/upstream/mode?value=down'
upstream mode = DOWN
HTTP 503
{
"status": "DOWN",
"externalApi": {
"details": {
"error": "org.springframework.web.client.HttpServerErrorException$ServiceUnavailable: 503 : \"upstream unavailable\"",
"url": "http://localhost:8080/stub/upstream/ping",
"latencyMs": 67,
"timeoutMs": 750
},
"status": "DOWN"
}
}
The aggregate went DOWN and /actuator/health now answers 503. If that URL is your
Kubernetes readiness probe, every pod in the deployment has just left the load
balancer because a third party had a bad minute.
--- a single component, addressed directly ---
$ curl -s http://localhost:8080/actuator/health/ordersDatabase
{
"details": {
"orders": 3,
"queryMs": 0,
"slowThresholdMs": 250
},
"status": "UP"
}
$ curl -s http://localhost:8080/actuator/health/kafka
{
"details": {
"bootstrap": "localhost:9092",
"tuned": true,
"error": "TimeoutException: null",
"probeMs": 1500,
"budgetMs": 1500
},
"status": "DOWN"
}
--- restore ---
upstream mode = UP

View File

@@ -0,0 +1,65 @@
### profile: groups
--- baseline: upstream UP ---
GET /actuator/health HTTP 503
GET /actuator/health/liveness HTTP 200
GET /actuator/health/readiness HTTP 200
GET /actuator/health/startup HTTP 200
--- upstream DOWN (a third party is having an outage) ---
GET /actuator/health HTTP 503
GET /actuator/health/liveness HTTP 200
GET /actuator/health/readiness HTTP 503
GET /actuator/health/startup HTTP 200
liveness stayed 200. readiness went 503.
Kubernetes takes this instance out of the Service and leaves the process alone.
Wire readiness to /actuator/health and you get a restart loop instead.
--- what each group contains ---
$ curl -s http://localhost:8080/actuator/health/liveness
{
"components": {
"diskSpace": {
"details": {
"total": 10213466112,
"free": 3877605376,
"threshold": 10485760,
"path": "/tmp/work/spring-boot-demo/.",
"exists": true
},
"status": "UP"
},
"livenessState": {
"status": "UP"
}
},
"status": "UP"
}
$ curl -s http://localhost:8080/actuator/health/readiness
{
"components": {
"externalApi": {
"details": {
"error": "org.springframework.web.client.HttpServerErrorException$ServiceUnavailable: 503 : \"upstream unavailable\"",
"url": "http://localhost:8080/stub/upstream/ping",
"latencyMs": 66,
"timeoutMs": 750
},
"status": "DOWN"
},
"ordersDatabase": {
"details": {
"orders": 3,
"queryMs": 0,
"slowThresholdMs": 250
},
"status": "UP"
},
"readinessState": {
"status": "UP"
}
},
"status": "DOWN"
}

View File

@@ -0,0 +1,15 @@
### Kafka health check against an unreachable broker -- naive
$ time curl -s http://localhost:8080/actuator/health/kafka
{
"details": {
"bootstrap": "localhost:9092",
"tuned": false,
"error": "TimeoutException: Timed out waiting for a node assignment. Call: listNodes",
"probeMs": 60002,
"budgetMs": 180000
},
"status": "DOWN"
}
wall clock: 60.2 s

View File

@@ -0,0 +1,15 @@
### Kafka health check against an unreachable broker -- tuned
$ time curl -s http://localhost:8080/actuator/health/kafka
{
"details": {
"bootstrap": "localhost:9092",
"tuned": true,
"error": "TimeoutException: null",
"probeMs": 1500,
"budgetMs": 1500
},
"status": "DOWN"
}
wall clock: 1.6 s

View File

@@ -0,0 +1,18 @@
### profile: details upstream deliberately sleeping 30s per request
### demo.upstream.timeout-ms = 750, so the indicator gives up long before the stub replies
$ time curl -s http://localhost:8080/actuator/health/externalApi
{
"details": {
"error": "org.springframework.web.client.ResourceAccessException: I/O error on GET request for \"http://localhost:8080/stub/upstream/ping\": Read timed out",
"url": "http://localhost:8080/stub/upstream/ping",
"latencyMs": 753,
"timeoutMs": 750
},
"status": "DOWN"
}
wall clock: 0.84 s
The read timeout is what bounds this, not the endpoint. Remove setReadTimeout from
ExternalApiHealthIndicator and this call blocks for the full 30 seconds, holding a
Tomcat worker the whole time.

View File

@@ -0,0 +1,83 @@
<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 https://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<parent>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-parent</artifactId>
<version>4.1.1</version>
<relativePath/>
</parent>
<groupId>com.ankurm</groupId>
<artifactId>actuator-production</artifactId>
<version>1.0.0</version>
<name>actuator-production</name>
<description>Spring Boot Actuator in production: endpoints, security, custom health indicators</description>
<properties>
<java.version>25</java.version>
</properties>
<dependencies>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-webmvc</artifactId>
</dependency>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-actuator</artifactId>
</dependency>
<!-- Spring Boot 4 split RestClient out of the web starter. Without this, injecting
RestClient.Builder fails at startup with "required a bean of type
'org.springframework.web.client.RestClient$Builder' that could not be found".
See docs/02-boot-4-changes.md. -->
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-restclient</artifactId>
</dependency>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-security</artifactId>
</dependency>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-data-jpa</artifactId>
</dependency>
<dependency>
<groupId>com.h2database</groupId>
<artifactId>h2</artifactId>
<scope>runtime</scope>
</dependency>
<dependency>
<groupId>org.apache.kafka</groupId>
<artifactId>kafka-clients</artifactId>
</dependency>
<dependency>
<groupId>io.micrometer</groupId>
<artifactId>micrometer-registry-prometheus</artifactId>
<scope>runtime</scope>
</dependency>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-test</artifactId>
<scope>test</scope>
</dependency>
<dependency>
<groupId>org.springframework.security</groupId>
<artifactId>spring-security-test</artifactId>
<scope>test</scope>
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-maven-plugin</artifactId>
</plugin>
</plugins>
</build>
</project>

View File

@@ -0,0 +1,40 @@
#!/usr/bin/env bash
# The three custom indicators, and what happens when the upstream goes down.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/07-custom-health-indicators.txt
B=http://localhost:8080
{
echo "### profile: details (show-details: always, show-components: always)"
echo
echo "--- upstream UP ---"
curl -s -o /dev/null -u ops:ops-password -X POST "$B/stub/upstream/mode?value=up"
sleep 1
echo "\$ curl -s -o /dev/null -w '%{http_code}' $B/actuator/health"
curl -s -o /dev/null -w ' HTTP %{http_code}\n' -u ops:ops-password "$B/actuator/health"
curl -s -u ops:ops-password "$B/actuator/health" | python3 -m json.tool
echo
echo "--- flip the upstream to DOWN, change nothing else ---"
echo "\$ curl -s -X POST '$B/stub/upstream/mode?value=down'"
curl -s -u ops:ops-password -X POST "$B/stub/upstream/mode?value=down"; echo
sleep 1
curl -s -o /dev/null -w ' HTTP %{http_code}\n' -u ops:ops-password "$B/actuator/health"
curl -s -u ops:ops-password "$B/actuator/health" | python3 -c '
import json,sys
d=json.load(sys.stdin)
print(json.dumps({"status":d["status"],"externalApi":d["components"]["externalApi"]},indent=2))'
echo
echo " The aggregate went DOWN and /actuator/health now answers 503. If that URL is your"
echo " Kubernetes readiness probe, every pod in the deployment has just left the load"
echo " balancer because a third party had a bad minute."
echo
echo "--- a single component, addressed directly ---"
echo "\$ curl -s $B/actuator/health/ordersDatabase"
curl -s -u ops:ops-password "$B/actuator/health/ordersDatabase" | python3 -m json.tool
echo "\$ curl -s $B/actuator/health/kafka"
curl -s -u ops:ops-password "$B/actuator/health/kafka" | python3 -m json.tool
echo
echo "--- restore ---"
curl -s -u ops:ops-password -X POST "$B/stub/upstream/mode?value=up"; echo
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,22 @@
#!/usr/bin/env bash
# What Actuator exposes when you add the starter and configure nothing.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/01-default-exposure.txt
{
echo "### Spring Boot Actuator, starter added, ZERO management.* configuration"
echo
echo "\$ curl -s -u ops:ops-password http://localhost:8080/actuator"
curl -s -u ops:ops-password http://localhost:8080/actuator | python3 -m json.tool
echo
echo "\$ curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password http://localhost:8080/actuator/env"
curl -s -o /dev/null -w '%{http_code}\n' -u ops:ops-password http://localhost:8080/actuator/env
echo " 404 = discovered but NOT exposed over HTTP. Exposure and existence are different things."
echo
echo "\$ curl -s -u ops:ops-password http://localhost:8080/actuator/health"
curl -s -u ops:ops-password http://localhost:8080/actuator/health | python3 -m json.tool
echo
echo " Only 'health' is web-exposed by default. show-details defaults to 'never', so even an"
echo " authenticated caller sees a bare status until you say otherwise."
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,13 @@
#!/usr/bin/env bash
# The authoritative endpoint list: what the RUNNING application publishes with exposure = "*".
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/02-endpoint-catalogue.txt
{
echo "### Every web-exposed endpoint, read from the running application"
echo "### profiles: exposeall,open management.endpoints.web.exposure.include: \"*\""
echo
echo "\$ curl -s http://localhost:8080/actuator/diag"
curl -s http://localhost:8080/actuator/diag | python3 -m json.tool
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,36 @@
#!/usr/bin/env bash
# Health groups: keeping a dependency outage out of the liveness probe.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/08-groups-and-probes.txt
B=http://localhost:8080/actuator/health
{
echo "### profile: groups"
echo
echo "--- baseline: upstream UP ---"
for g in "" liveness readiness startup; do
u="$B${g:+/$g}"
printf ' %-40s HTTP %s\n' "GET ${u#http://localhost:8080}" "$(curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password "$u")"
done
echo
echo "--- upstream DOWN (a third party is having an outage) ---"
curl -s -o /dev/null -u ops:ops-password -X POST "http://localhost:8080/stub/upstream/mode?value=down"
sleep 1
for g in "" liveness readiness startup; do
u="$B${g:+/$g}"
printf ' %-40s HTTP %s\n' "GET ${u#http://localhost:8080}" "$(curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password "$u")"
done
echo
echo " liveness stayed 200. readiness went 503."
echo " Kubernetes takes this instance out of the Service and leaves the process alone."
echo " Wire readiness to /actuator/health and you get a restart loop instead."
echo
echo "--- what each group contains ---"
echo "\$ curl -s $B/liveness"
curl -s -u ops:ops-password "$B/liveness" | python3 -m json.tool
echo "\$ curl -s $B/readiness"
curl -s -u ops:ops-password "$B/readiness" | python3 -m json.tool
echo
curl -s -o /dev/null -u ops:ops-password -X POST "http://localhost:8080/stub/upstream/mode?value=up"
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,23 @@
#!/usr/bin/env bash
# The endpoint that really does hand over your secrets - once you turn it on.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/04-heapdump-leak.txt
B=http://localhost:8080/actuator
{
echo "### profiles: exposeall,open PLUS --management.endpoint.heapdump.access=unrestricted"
echo
echo "\$ curl -s -o /tmp/heap.hprof -w 'status=%{http_code} bytes=%{size_download} type=%{content_type}' $B/heapdump"
curl -s -o /tmp/heap.hprof -w 'status=%{http_code} bytes=%{size_download} type=%{content_type}\n' "$B/heapdump"
echo
echo "\$ strings /tmp/heap.hprof | grep -c 'S3CRET-partner-credential'"
strings /tmp/heap.hprof 2>/dev/null | grep -c 'S3CRET-partner-credential'
echo "\$ strings /tmp/heap.hprof | grep -o 'not-a-real-password[^\"]*' | head -1"
strings /tmp/heap.hprof 2>/dev/null | grep -o 'not-a-real-password[^\"]*' | head -1
echo
echo " /actuator/env masked both of these to ******."
echo " /actuator/heapdump handed over the process memory that contains them in plaintext."
echo " Sanitisation is a property-rendering feature. It is not a security boundary."
rm -f /tmp/heap.hprof
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,19 @@
#!/usr/bin/env bash
# The single most useful number in this repository: how long a naive Kafka health check blocks.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/09-kafka-timeout.txt
MODE="${1:-tuned}"
B=http://localhost:8080/actuator/health/kafka
{
echo "### Kafka health check against an unreachable broker -- ${MODE}"
echo
echo "\$ time curl -s $B"
s=$(date +%s.%N)
body=$(curl -s -u ops:ops-password --max-time 180 "$B")
e=$(date +%s.%N)
echo "$body" | python3 -m json.tool 2>/dev/null || echo "$body"
echo
printf 'wall clock: %.1f s\n' "$(echo "$e - $s" | bc)"
} > "$OUT.$MODE" 2>&1
echo "wrote $OUT.$MODE"

View File

@@ -0,0 +1,34 @@
#!/usr/bin/env bash
# Actuator on its own port and path.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/06-management-port.txt
{
echo "### profile: mgmtport"
echo "### management.server.port: 9001 / management.server.address: 127.0.0.1 / base-path: /manage"
echo
echo "--- the application port no longer serves Actuator at all ---"
printf ' %-52s %s\n' "GET :8080/actuator/health" "$(curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:8080/actuator/health)"
printf ' %-52s %s\n' "GET :8080/manage/health" "$(curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:8080/manage/health)"
printf ' %-52s %s\n' "GET :8080/orders/count" "$(curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password http://127.0.0.1:8080/orders/count)"
echo
echo "--- the management port serves it on the new base path ---"
printf ' %-52s %s\n' "GET :9001/manage/health" "$(curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password http://127.0.0.1:9001/manage/health)"
printf ' %-52s %s\n' "GET :9001/actuator/health" "$(curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password http://127.0.0.1:9001/actuator/health)"
printf ' %-52s %s\n' "GET :9001/orders/count" "$(curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password http://127.0.0.1:9001/orders/count)"
echo
echo " Note the last line. The management context has its own DispatcherServlet and does NOT"
echo " see application controllers. That is the isolation you are paying for."
echo
echo "--- what the management context reports about itself ---"
echo "\$ curl -s -u ops:ops-password http://127.0.0.1:9001/manage/diag"
curl -s -u ops:ops-password http://127.0.0.1:9001/manage/diag | python3 -m json.tool
echo
echo "--- listening sockets ---"
echo "\$ ss -ltn | grep -E ':(8080|9001)'"
ss -ltn 2>/dev/null | grep -E ':(8080|9001)' || netstat -ltn 2>/dev/null | grep -E ':(8080|9001)'
echo
echo " 9001 is bound to 127.0.0.1 only. 8080 is bound to *. An ingress that forwards to 8080"
echo " cannot reach Actuator no matter how the security rules are written."
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,53 @@
#!/usr/bin/env bash
# What an unauthenticated caller actually gets when exposure is "*" and the chain permits all.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/03-open-actuator-leak.txt
B=http://localhost:8080/actuator
{
echo "### profiles: exposeall,open -- NO credentials are sent on any request below"
echo
echo "--- 1. /actuator/env does NOT leak values in Spring Boot 4 ---"
echo "\$ curl -s $B/env/spring.datasource.password | jq .property"
curl -s "$B/env/spring.datasource.password" | python3 -c 'import json,sys;print(json.dumps(json.load(sys.stdin)["property"],indent=2))'
echo "\$ curl -s $B/env/acme.partner.credential | jq .property"
curl -s "$B/env/acme.partner.credential" | python3 -c 'import json,sys;print(json.dumps(json.load(sys.stdin)["property"],indent=2))'
echo
echo " Note the second one. 'acme.partner.credential' matches none of the classic"
echo " password/secret/token key patterns, and it is still masked. Masking is driven by"
echo " management.endpoint.env.show-values, which defaults to 'never' - not by key names."
echo
echo "--- 2. /actuator/heapdump is NOT exposed by 'include: \"*\"' ---"
echo "\$ curl -s -o /dev/null -w '%{http_code}' $B/heapdump"
curl -s -o /dev/null -w '%{http_code}\n' "$B/heapdump"
echo " management.endpoint.heapdump.access defaults to 'none'. So does shutdown."
echo " They are the only two endpoints that do."
echo
echo "--- 3. /actuator/loggers: an unauthenticated WRITE ---"
echo "\$ curl -s -X POST -d '{\"configuredLevel\":\"TRACE\"}' -H 'Content-Type: application/json' $B/loggers/org.springframework"
curl -s -o /dev/null -w ' status=%{http_code}\n' -X POST -H 'Content-Type: application/json' \
-d '{"configuredLevel":"TRACE"}' "$B/loggers/org.springframework"
echo "\$ curl -s $B/loggers/org.springframework"
curl -s "$B/loggers/org.springframework" | python3 -m json.tool
curl -s -o /dev/null -X POST -H 'Content-Type: application/json' \
-d '{"configuredLevel":null}' "$B/loggers/org.springframework"
echo " (level reset). An attacker who can flip your root logger to TRACE has both a"
echo " denial-of-service primitive and a way to get request bodies written to disk."
echo
echo "--- 4. /actuator/beans and /actuator/mappings: your whole application, described ---"
echo "\$ curl -s $B/mappings | python3 -c 'count the URL patterns'"
curl -s "$B/mappings" | python3 -c '
import json,sys
d=json.load(sys.stdin)
n=0
for ctx in d["contexts"].values():
for m in ctx["mappings"].get("dispatcherServlets",{}).values():
n+=len(m)
print(f" {n} servlet mappings disclosed")'
curl -s "$B/beans" | python3 -c '
import json,sys
d=json.load(sys.stdin)
n=sum(len(c["beans"]) for c in d["contexts"].values())
print(f" {n} beans disclosed, each with its type and dependencies")'
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,45 @@
#!/usr/bin/env bash
# The authorisation matrix produced by SecuredActuatorConfig.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/05-secured-matrix.txt
B=http://localhost:8080/actuator
probe() { # $1 label, $2 path, $3 curl auth args...
local label="$1" path="$2"; shift 2
printf ' %-46s %s\n' "$label" "$(curl -s -o /dev/null -w '%{http_code}' "$@" "$B/$path")"
}
{
echo "### profile: secured (SecuredActuatorConfig + application-secured.yaml)"
echo "### exposure is \"*\" - the security chain, not the exposure list, is what protects it"
echo
echo "ANONYMOUS"
probe "GET /actuator/health" health
probe "GET /actuator/info" info
probe "GET /actuator/env" env
probe "GET /actuator/beans" beans
probe "GET /actuator/threaddump" threaddump
probe "GET /actuator (the links index)" ""
echo
echo "AUTHENTICATED as ops (ROLE_ACTUATOR)"
probe "GET /actuator/health" health -u ops:ops-password
probe "GET /actuator/env" env -u ops:ops-password
probe "GET /actuator/beans" beans -u ops:ops-password
probe "GET /actuator/threaddump" threaddump -u ops:ops-password
echo
echo "WRONG PASSWORD"
probe "GET /actuator/env" env -u ops:wrong
echo
echo "--- health body, anonymous (show-details: when-authorized) ---"
curl -s "$B/health" | python3 -m json.tool
echo
echo "--- health body, authenticated as ROLE_ACTUATOR ---"
curl -s -u ops:ops-password "$B/health" | python3 -m json.tool
echo
echo " Same endpoint, same status code, different body. An anonymous prober learns that the"
echo " service is unhealthy but not WHICH dependency is unhealthy."
echo
echo "--- the business endpoint is untouched by the actuator chain ---"
printf ' %-46s %s\n' "GET /orders/count (anonymous)" \
"$(curl -s -o /dev/null -w '%{http_code}' http://localhost:8080/orders/count)"
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,24 @@
#!/usr/bin/env bash
# A health indicator with no timeout, and what it does to the endpoint.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/10-slow-upstream.txt
B=http://localhost:8080
{
echo "### profile: details upstream deliberately sleeping 30s per request"
echo "### demo.upstream.timeout-ms = 750, so the indicator gives up long before the stub replies"
echo
curl -s -o /dev/null -u ops:ops-password -X POST "$B/stub/upstream/mode?value=slow"
sleep 1
echo "\$ time curl -s $B/actuator/health/externalApi"
s=$(date +%s.%N)
curl -s -u ops:ops-password --max-time 60 "$B/actuator/health/externalApi" | python3 -m json.tool
e=$(date +%s.%N)
printf 'wall clock: %.2f s\n' "$(echo "$e - $s" | bc)"
echo
echo " The read timeout is what bounds this, not the endpoint. Remove setReadTimeout from"
echo " ExternalApiHealthIndicator and this call blocks for the full 30 seconds, holding a"
echo " Tomcat worker the whole time."
curl -s -o /dev/null -u ops:ops-password -X POST "$B/stub/upstream/mode?value=up"
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,19 @@
#!/usr/bin/env bash
# Exactly what this repository was built and run against.
set -uo pipefail
cd "$(dirname "$0")/.."
OUT=docs/output/00-versions.txt
{
echo "\$ java -version"
java -version 2>&1 | grep -viE 'JAVA_TOOL_OPTIONS|Picked up'
echo
echo "\$ mvn -v | head -3"
mvn -v 2>&1 | grep -viE 'WARNING|sun\.misc|Picked up|Please consider' | head -3
echo
echo "\$ mvn dependency:list -- the resolved versions behind spring-boot-starter-parent 4.1.1"
mvn -B dependency:list 2>/dev/null \
| grep -oE '(org\.springframework[a-z.]*|io\.micrometer|org\.apache\.kafka|com\.h2database):[A-Za-z0-9.-]+:jar:[0-9][A-Za-z0-9.-]*' \
| awk -F':' '{printf " %-52s %s\n", $1":"$2, $4}' | sort -u \
| grep -E 'spring-boot-actuator |spring-boot-health |spring-boot |spring-core |spring-web |spring-security-core |spring-security-web |micrometer-core |micrometer-registry-prometheus |kafka-clients |h2 |spring-boot-restclient '
} > "$OUT" 2>&1
echo "wrote $OUT"

View File

@@ -0,0 +1,8 @@
#!/usr/bin/env bash
# Shared environment. Point JAVA_HOME at a JDK 25 (or newer) installation.
: "${JAVA_HOME:?set JAVA_HOME to a JDK 25+ installation}"
export PATH="$JAVA_HOME/bin:$PATH"
MVN="${MVN:-mvn}"
APP_MAIN="com.ankurm.actuator.ActuatorProductionApplication"
APP_PORT="${APP_PORT:-8080}"
MGMT_PORT="${MGMT_PORT:-9001}"

View File

@@ -0,0 +1,70 @@
#!/usr/bin/env bash
# Regenerate every file in docs/output/ from scratch.
#
# JAVA_HOME=/path/to/jdk25 ./scripts/run-all.sh
#
# Each scenario starts a fresh JVM, runs its probes, and stops it. Timing figures vary between
# machines; the status codes and bodies do not.
set -uo pipefail
cd "$(dirname "$0")/.."
: "${JAVA_HOME:?set JAVA_HOME to a JDK 25+ installation}"
export PATH="$JAVA_HOME/bin:$PATH"
JAR=target/actuator-production-1.0.0.jar
[ -f "$JAR" ] || mvn -B -q -DskipTests package
HEAP="${HEAP:--Xmx384m}"
APP_PID=""
start() { # $1 profiles, rest: extra --args
local profiles="$1"; shift
java $HEAP -jar "$JAR" --spring.profiles.active="$profiles" "$@" > /tmp/actuator-demo.log 2>&1 &
APP_PID=$!
local port=8080 url="http://127.0.0.1:8080/orders/count"
for _ in $(seq 1 90); do
c=$(curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password "$url" || true)
[ -n "$c" ] && [ "$c" != "000" ] && return 0
sleep 1
done
echo "FAILED TO START (profiles=$profiles)" >&2; tail -30 /tmp/actuator-demo.log >&2; return 1
}
stop() {
[ -n "$APP_PID" ] && kill -9 "$APP_PID" 2>/dev/null
wait "$APP_PID" 2>/dev/null
APP_PID=""
for _ in $(seq 1 40); do
(exec 3<>/dev/tcp/127.0.0.1/8080) 2>/dev/null || break
sleep 0.25
done
exec 3<&- 2>/dev/null || true
}
trap stop EXIT
mkdir -p docs/output
scripts/demo-versions.sh
start "" && { scripts/demo-default-exposure.sh; stop; }
start "exposeall,open" && { scripts/demo-endpoint-catalogue.sh;
scripts/demo-open-actuator.sh; stop; }
start "exposeall,open" --management.endpoint.heapdump.access=unrestricted \
&& { scripts/demo-heapdump-leak.sh; stop; }
start "secured" && { scripts/demo-secured.sh; stop; }
start "details" && { scripts/demo-custom-health.sh;
scripts/demo-slow-indicator.sh;
scripts/demo-kafka-timeout.sh tuned; stop; }
start "details,kafkanaive" && { scripts/demo-kafka-timeout.sh naive; stop; }
start "groups" && { scripts/demo-groups-probes.sh; stop; }
# The management-port scenario readies on a different URL, so it is started by hand.
java $HEAP -jar "$JAR" --spring.profiles.active=mgmtport > /tmp/actuator-demo.log 2>&1 &
APP_PID=$!
for _ in $(seq 1 90); do
c=$(curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password http://127.0.0.1:9001/manage/health || true)
[ -n "$c" ] && [ "$c" != "000" ] && break
sleep 1
done
scripts/demo-management-port.sh
stop
echo
echo "docs/output/ regenerated:"
ls -1 docs/output/

View File

@@ -0,0 +1,26 @@
#!/usr/bin/env bash
# Start the application with the given comma-separated profiles and block until it answers.
# ./scripts/run.sh # defaults
# ./scripts/run.sh exposeall,open # every endpoint, no authentication
set -euo pipefail
cd "$(dirname "$0")/.."
source scripts/env.sh
PROFILES="${1:-}"
LOG="${LOG:-/tmp/actuator-demo.log}"
scripts/stop.sh
ARGS=(-B -o org.springframework.boot:spring-boot-maven-plugin:run)
[ -n "$PROFILES" ] && ARGS+=("-Dspring-boot.run.profiles=$PROFILES")
setsid nohup "$MVN" "${ARGS[@]}" > "$LOG" 2>&1 < /dev/null &
READY_URL="${READY_URL:-http://127.0.0.1:${APP_PORT}/orders/count}"
for _ in $(seq 1 90); do
code=$(curl -s -o /dev/null -w '%{http_code}' -u ops:ops-password "$READY_URL" || true)
[ "$code" != "000" ] && [ -n "$code" ] && exit 0
sleep 2
done
echo "application did not start; tail of $LOG:" >&2
tail -40 "$LOG" >&2
exit 1

View File

@@ -0,0 +1,14 @@
#!/usr/bin/env bash
# Kill by main class, never by a pattern that could match this script's own command line.
# `pkill -f spring-boot` matches the shell running it and takes the terminal with it.
set -u
for p in $(ps -eo pid,args | grep '[A]ctuatorProductionApplication' | awk '{print $1}'); do
kill -9 "$p" 2>/dev/null || true
done
# Wait for the port to actually close. Killing the PID is not the same as the socket being
# free, and a stale listener looks exactly like your configuration change having no effect.
for _ in $(seq 1 40); do
if ! (exec 3<>/dev/tcp/127.0.0.1/"${APP_PORT:-8080}") 2>/dev/null; then break; fi
sleep 0.25
done
exec 3<&- 2>/dev/null || true

View File

@@ -0,0 +1,11 @@
package com.ankurm.actuator;
import org.springframework.boot.SpringApplication;
import org.springframework.boot.autoconfigure.SpringBootApplication;
@SpringBootApplication
public class ActuatorProductionApplication {
public static void main(String[] args) {
SpringApplication.run(ActuatorProductionApplication.class, args);
}
}

View File

@@ -0,0 +1,37 @@
package com.ankurm.actuator.config;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
import org.springframework.context.annotation.Profile;
import org.springframework.security.config.annotation.web.builders.HttpSecurity;
import org.springframework.security.config.http.SessionCreationPolicy;
import org.springframework.security.web.SecurityFilterChain;
/**
* The baseline chain used by every profile that does not bring its own.
*
* <p>It exists for one reason worth knowing about: Spring Boot's auto-configured security is
* browser-shaped. It turns CSRF protection on, which means an authenticated {@code POST} from
* curl is rejected before it reaches your controller. While building this repository that
* showed up as a {@code 401} on {@code POST /stub/upstream/mode} while {@code GET /orders/count}
* with the same credentials returned {@code 200} &mdash; a confusing pair of results that has
* nothing to do with the credentials.
*
* <p>The same trap catches people calling {@code POST /actuator/loggers/{name}} from a script.
* Actuator is a machine-to-machine API; give it a stateless, CSRF-free chain.
*
* <p>See docs/04-securing-actuator.md.
*/
@Configuration
@Profile("!open & !secured")
public class DemoSecurityConfig {
@Bean
SecurityFilterChain defaultChain(HttpSecurity http) throws Exception {
http.authorizeHttpRequests((requests) -> requests.anyRequest().authenticated())
.httpBasic((basic) -> { })
.csrf((csrf) -> csrf.disable())
.sessionManagement((s) -> s.sessionCreationPolicy(SessionCreationPolicy.STATELESS));
return http.build();
}
}

View File

@@ -0,0 +1,33 @@
package com.ankurm.actuator.config;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
import org.springframework.context.annotation.Profile;
import org.springframework.security.config.annotation.web.builders.HttpSecurity;
import org.springframework.security.web.SecurityFilterChain;
/**
* The misconfiguration, kept in the repository on purpose so the damage can be measured rather
* than asserted.
*
* <p>This is not a strawman. It is what you get when somebody adds Spring Security, finds that
* the generated password broke their smoke tests, and reaches for the shortest fix that makes
* the tests pass again. Combined with
* {@code management.endpoints.web.exposure.include: "*"} it publishes {@code /actuator/env},
* {@code /actuator/heapdump} and {@code /actuator/shutdown} to anyone who can reach the port.
*
* <p>Run {@code ./scripts/demo-open-actuator.sh} to see exactly what leaks.
*
* <p>See docs/04-securing-actuator.md.
*/
@Configuration
@Profile("open")
public class OpenActuatorConfig {
@Bean
SecurityFilterChain permitEverything(HttpSecurity http) throws Exception {
http.authorizeHttpRequests((requests) -> requests.anyRequest().permitAll())
.csrf((csrf) -> csrf.disable());
return http.build();
}
}

View File

@@ -0,0 +1,66 @@
package com.ankurm.actuator.config;
import org.springframework.boot.security.autoconfigure.actuate.web.servlet.EndpointRequest;
import org.springframework.boot.health.actuate.endpoint.HealthEndpoint;
import org.springframework.boot.actuate.info.InfoEndpoint;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
import org.springframework.context.annotation.Profile;
import org.springframework.core.annotation.Order;
import org.springframework.security.config.annotation.web.builders.HttpSecurity;
import org.springframework.security.web.SecurityFilterChain;
/**
* The configuration you actually want in production.
*
* <p>Three things matter here and each of them is a mistake people make:
*
* <ol>
* <li><strong>Match on {@code EndpointRequest.toAnyEndpoint()}, not on a path string.</strong>
* A rule written as {@code "/actuator/**"} silently stops matching the moment somebody
* sets {@code management.endpoints.web.base-path}, or moves Actuator to its own port. The
* matcher asks the endpoint registry, so it follows the configuration.
* <li><strong>This chain is ordered ahead of the application's chain.</strong> Without an
* explicit order, whichever chain Spring happens to register first wins for a given
* request, and the application chain's {@code permitAll} can swallow the Actuator paths.
* <li><strong>Only {@code health} and {@code info} are anonymous</strong>, and even health is
* details-free for anonymous callers &mdash; see {@code show-details: when-authorized} in
* application-secured.yaml. Everything else needs the ACTUATOR role.
* </ol>
*
* <p>In Spring Boot 4 {@code EndpointRequest} moved to
* {@code org.springframework.boot.security.autoconfigure.actuate.web.servlet}; in 3.x it was
* {@code org.springframework.boot.actuate.autoconfigure.security.servlet}. Same class, same
* methods, new package. See docs/02-boot-4-changes.md.
*
* <p>See docs/04-securing-actuator.md.
*/
@Configuration
@Profile("secured")
public class SecuredActuatorConfig {
@Bean
@Order(1)
SecurityFilterChain actuatorChain(HttpSecurity http) throws Exception {
http.securityMatcher(EndpointRequest.toAnyEndpoint())
.authorizeHttpRequests((requests) -> requests
.requestMatchers(EndpointRequest.to(HealthEndpoint.class, InfoEndpoint.class)).permitAll()
.anyRequest().hasRole("ACTUATOR"))
.httpBasic((basic) -> { })
// Actuator is a machine-to-machine API. Sessions and CSRF tokens are for
// browsers; a stateless chain avoids the 403-on-POST that catches people
// calling /actuator/loggers or /actuator/shutdown from curl.
.csrf((csrf) -> csrf.disable())
.sessionManagement((session) -> session.sessionCreationPolicy(
org.springframework.security.config.http.SessionCreationPolicy.STATELESS));
return http.build();
}
@Bean
@Order(2)
SecurityFilterChain applicationChain(HttpSecurity http) throws Exception {
http.authorizeHttpRequests((requests) -> requests.anyRequest().permitAll())
.csrf((csrf) -> csrf.disable());
return http.build();
}
}

View File

@@ -0,0 +1,80 @@
package com.ankurm.actuator.health;
import java.time.Duration;
import java.time.Instant;
import org.springframework.beans.factory.annotation.Value;
import org.springframework.boot.health.contributor.AbstractHealthIndicator;
import org.springframework.boot.health.contributor.Health;
import org.springframework.stereotype.Component;
import org.springframework.web.client.RestClient;
/**
* Probes a third-party HTTP dependency.
*
* <p>Two things here are the whole point of the chapter:
*
* <ul>
* <li>The client carries explicit connect and read timeouts. A health indicator without a
* timeout inherits the JVM default, which is "wait forever" &mdash; and a hung indicator
* hangs {@code /actuator/health}, which hangs the load balancer probe, which takes the
* whole fleet out. See docs/06-health-indicator-failure-modes.md.
* <li>It is registered under the group {@code readiness} only, never {@code liveness}. A
* third party being down must not restart your pod. See docs/07-groups-and-probes.md.
* </ul>
*
* <p>In Spring Boot 4 the base class moved: {@code AbstractHealthIndicator} and {@code Health}
* are in {@code org.springframework.boot.health.contributor}, not
* {@code org.springframework.boot.actuate.health}. See docs/02-boot-4-changes.md.
*/
@Component("externalApi")
public class ExternalApiHealthIndicator extends AbstractHealthIndicator {
private final RestClient client;
private final String url;
private final Duration budget;
public ExternalApiHealthIndicator(
RestClient.Builder builder,
@Value("${demo.upstream.url:http://localhost:8080/stub/upstream/ping}") String url,
@Value("${demo.upstream.timeout-ms:750}") long timeoutMs,
@Value("${demo.upstream.username:}") String username,
@Value("${demo.upstream.password:}") String password) {
this.url = url;
this.budget = Duration.ofMillis(timeoutMs);
var factory = new org.springframework.http.client.SimpleClientHttpRequestFactory();
factory.setConnectTimeout(this.budget);
factory.setReadTimeout(this.budget);
if (!username.isEmpty()) {
// Real upstreams are authenticated, and a health indicator that forgets its
// credentials reports DOWN for a reason that has nothing to do with the upstream.
// This repository hit exactly that: the first captured run showed
// "error": "HttpClientErrorException$Unauthorized: 401"
// which looks like an outage and was actually a missing Authorization header.
builder = builder.defaultHeaders((headers) -> headers.setBasicAuth(username, password));
}
this.client = builder.requestFactory(factory).build();
}
@Override
protected void doHealthCheck(Health.Builder builder) {
Instant start = Instant.now();
try {
String body = this.client.get().uri(this.url).retrieve().body(String.class);
builder.up()
.withDetail("url", this.url)
.withDetail("response", body)
.withDetail("latencyMs", Duration.between(start, Instant.now()).toMillis())
.withDetail("timeoutMs", this.budget.toMillis());
}
catch (Exception ex) {
// down(ex) records the exception under the "error" detail key. That detail is only
// rendered when show-details permits it, which is why an unauthenticated caller
// still sees a bare {"status":"DOWN"}.
builder.down(ex)
.withDetail("url", this.url)
.withDetail("latencyMs", Duration.between(start, Instant.now()).toMillis())
.withDetail("timeoutMs", this.budget.toMillis());
}
}
}

View File

@@ -0,0 +1,121 @@
package com.ankurm.actuator.health;
import java.time.Duration;
import java.time.Instant;
import java.util.HashMap;
import java.util.Map;
import java.util.concurrent.TimeUnit;
import org.apache.kafka.clients.CommonClientConfigs;
import org.apache.kafka.clients.admin.AdminClient;
import org.apache.kafka.clients.admin.AdminClientConfig;
import org.apache.kafka.clients.admin.DescribeClusterOptions;
import org.apache.kafka.clients.admin.DescribeClusterResult;
import org.springframework.beans.factory.annotation.Value;
import org.springframework.boot.health.contributor.AbstractHealthIndicator;
import org.springframework.boot.health.contributor.Health;
import org.springframework.stereotype.Component;
/**
* Checks a Kafka cluster with {@code AdminClient.describeCluster}, in two configurations, so
* the difference between them can be measured instead of argued about.
*
* <p><strong>Naive ({@code demo.kafka.tuned=false}).</strong> Create an {@code AdminClient}
* with only the bootstrap servers set and call {@code describeCluster().nodes().get()}. Against
* an unreachable broker that call returns after {@code default.api.timeout.ms}, whose Kafka
* default is <strong>60000</strong>. Not {@code request.timeout.ms} (30000) &mdash; that bounds
* one attempt, and {@code retries} defaults to {@code Integer.MAX_VALUE}, so attempts keep
* happening until the API timeout fires. A health endpoint that blocks for a minute is worse
* than no health endpoint: every probe piles another thread onto the container.
*
* <p><strong>Tuned ({@code demo.kafka.tuned=true}, the default here).</strong> Every one of the
* four bounds is set, and there is a hard {@code KafkaFuture.get(timeout)} outside them all.
* {@code metadata.recovery.strategy} is pinned to {@code none}: Kafka 4's default is
* {@code rebootstrap}, and a long-lived AdminClient pointed at a dead broker will otherwise
* fill your logs with rebootstrap lines from its background thread &mdash; this repository's
* first run produced 192 of them in a few seconds.
*
* <p>The {@code AdminClient} is created once and reused. Creating one per probe opens a fresh
* set of broker connections on every probe interval, which across a fleet is a denial of
* service against your own brokers.
*
* <p>No broker runs in this repository, so this indicator reports DOWN. That is deliberate:
* the transcripts in docs/output/ are real.
*
* <p>See docs/05-custom-health-indicators.md and docs/06-health-indicator-failure-modes.md.
*/
@Component("kafka")
public class KafkaHealthIndicator extends AbstractHealthIndicator implements AutoCloseable {
private final AdminClient admin;
private final String bootstrap;
private final Duration budget;
private final boolean tuned;
public KafkaHealthIndicator(
@Value("${demo.kafka.bootstrap:localhost:9092}") String bootstrap,
@Value("${demo.kafka.timeout-ms:1500}") long timeoutMs,
@Value("${demo.kafka.tuned:true}") boolean tuned) {
this.bootstrap = bootstrap;
this.budget = Duration.ofMillis(timeoutMs);
this.tuned = tuned;
Map<String, Object> config = new HashMap<>();
config.put(AdminClientConfig.BOOTSTRAP_SERVERS_CONFIG, bootstrap);
config.put(AdminClientConfig.CLIENT_ID_CONFIG, "health-" + (tuned ? "tuned" : "naive"));
if (tuned) {
// (1) how long one request may take
config.put(AdminClientConfig.REQUEST_TIMEOUT_MS_CONFIG, (int) timeoutMs);
// (2) how long the whole API call may take, retries included. Without this the
// call runs for 60s regardless of (1).
config.put(AdminClientConfig.DEFAULT_API_TIMEOUT_MS_CONFIG, (int) timeoutMs);
// (3) how long a TCP connect may take
config.put(AdminClientConfig.SOCKET_CONNECTION_SETUP_TIMEOUT_MS_CONFIG, (int) timeoutMs);
// (4) do not retry on the health path; the caller will probe again shortly
config.put(AdminClientConfig.RETRIES_CONFIG, 0);
config.put(CommonClientConfigs.RECONNECT_BACKOFF_MAX_MS_CONFIG, 5000);
config.put(CommonClientConfigs.METADATA_RECOVERY_STRATEGY_CONFIG, "none");
}
this.admin = AdminClient.create(config);
}
@Override
protected void doHealthCheck(Health.Builder builder) {
Instant start = Instant.now();
try {
// The naive path is deliberately the textbook one: no per-call timeout override,
// and a bare get(). Whatever bounds it is Kafka's own default, which is the point.
DescribeClusterResult result = this.tuned
? this.admin.describeCluster(
new DescribeClusterOptions().timeoutMs((int) this.budget.toMillis()))
: this.admin.describeCluster();
int nodes = this.tuned
? result.nodes().get(this.budget.toMillis(), TimeUnit.MILLISECONDS).size()
: result.nodes().get().size();
String clusterId = this.tuned
? result.clusterId().get(this.budget.toMillis(), TimeUnit.MILLISECONDS)
: result.clusterId().get();
builder.up()
.withDetail("bootstrap", this.bootstrap)
.withDetail("tuned", this.tuned)
.withDetail("clusterId", clusterId)
.withDetail("nodes", nodes)
.withDetail("probeMs", Duration.between(start, Instant.now()).toMillis());
}
catch (Exception ex) {
Throwable cause = (ex.getCause() != null) ? ex.getCause() : ex;
builder.down()
.withDetail("bootstrap", this.bootstrap)
.withDetail("tuned", this.tuned)
.withDetail("error", cause.getClass().getSimpleName() + ": " + cause.getMessage())
.withDetail("probeMs", Duration.between(start, Instant.now()).toMillis())
.withDetail("budgetMs", this.budget.toMillis());
}
}
@Override
public void close() {
this.admin.close(Duration.ofSeconds(2));
}
}

View File

@@ -0,0 +1,49 @@
package com.ankurm.actuator.health;
import java.time.Duration;
import java.time.Instant;
import org.springframework.boot.health.contributor.AbstractHealthIndicator;
import org.springframework.boot.health.contributor.Health;
import org.springframework.jdbc.core.JdbcTemplate;
import org.springframework.stereotype.Component;
/**
* A business-level database check, deliberately different from Boot's built-in
* {@code DataSourceHealthIndicator}.
*
* <p>The built-in one runs a validation query and reports UP if the connection works. That
* answers "can I reach the database", which is rarely the question that matters. This one
* answers "can I serve orders": it queries the table the application actually depends on, and
* it fails if the query is slow enough that requests would time out anyway.
*
* <p>Note the {@code @Component} name: it becomes the key in the {@code /actuator/health}
* response. Boot strips a trailing "HealthIndicator" from the bean name, so this bean could
* also have been named {@code ordersDatabaseHealthIndicator} for the same result.
*
* <p>See docs/05-custom-health-indicators.md.
*/
@Component("ordersDatabase")
public class OrdersDatabaseHealthIndicator extends AbstractHealthIndicator {
private final JdbcTemplate jdbc;
private final Duration slowThreshold = Duration.ofMillis(250);
public OrdersDatabaseHealthIndicator(JdbcTemplate jdbc) {
this.jdbc = jdbc;
// A query timeout is not optional. Without it the indicator blocks on the socket for
// as long as the driver's default allows, which on some drivers is indefinitely.
this.jdbc.setQueryTimeout(2);
}
@Override
protected void doHealthCheck(Health.Builder builder) {
Instant start = Instant.now();
Integer count = this.jdbc.queryForObject("SELECT COUNT(*) FROM orders", Integer.class);
Duration took = Duration.between(start, Instant.now());
builder.status(took.compareTo(this.slowThreshold) > 0 ? "DEGRADED" : "UP")
.withDetail("orders", count)
.withDetail("queryMs", took.toMillis())
.withDetail("slowThresholdMs", this.slowThreshold.toMillis());
}
}

View File

@@ -0,0 +1,26 @@
package com.ankurm.actuator.health;
import java.util.concurrent.atomic.AtomicReference;
import org.springframework.stereotype.Component;
/**
* Shared, mutable state for the stub upstream service. Flipping this at runtime is how the
* companion scripts make a health indicator go DOWN without needing a real outage.
*
* <p>See docs/05-custom-health-indicators.md.
*/
@Component
public class UpstreamState {
public enum Mode { UP, DOWN, SLOW }
private final AtomicReference<Mode> mode = new AtomicReference<>(Mode.UP);
public Mode get() {
return this.mode.get();
}
public void set(Mode mode) {
this.mode.set(mode);
}
}

View File

@@ -0,0 +1,96 @@
package com.ankurm.actuator.web;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.TreeMap;
import org.springframework.boot.actuate.endpoint.annotation.Endpoint;
import org.springframework.boot.actuate.endpoint.annotation.ReadOperation;
import org.springframework.boot.actuate.endpoint.web.ExposableWebEndpoint;
import org.springframework.boot.actuate.endpoint.web.WebEndpointsSupplier;
import org.springframework.boot.actuate.endpoint.web.WebOperation;
import org.springframework.boot.health.contributor.HealthContributor;
import org.springframework.boot.health.contributor.HealthContributors;
import org.springframework.boot.health.registry.HealthContributorRegistry;
import org.springframework.core.env.Environment;
import org.springframework.stereotype.Component;
/**
* Prints the runtime Actuator state that no configuration file will tell you: which endpoints
* were actually discovered and exposed on the web, which HTTP methods and paths each one
* publishes, and which health contributors are registered.
*
* <p>This exists because guessing at exposure from {@code application.yaml} is how people ship
* {@code /actuator/heapdump} to the internet. Ask the running application instead.
*
* <p><strong>Delete this before you ship.</strong> It is a debugging aid, and the list of
* exposed endpoints is itself reconnaissance.
*
* <p>See docs/08-diagnostics.md.
*/
@Component
@Endpoint(id = "diag")
public class DiagnosticsEndpoint {
private final WebEndpointsSupplier webEndpoints;
private final HealthContributorRegistry healthRegistry;
private final Environment environment;
public DiagnosticsEndpoint(WebEndpointsSupplier webEndpoints,
HealthContributorRegistry healthRegistry, Environment environment) {
this.webEndpoints = webEndpoints;
this.healthRegistry = healthRegistry;
this.environment = environment;
}
@ReadOperation
public Map<String, Object> diagnostics() {
Map<String, Object> result = new LinkedHashMap<>();
result.put("activeProfiles", List.of(this.environment.getActiveProfiles()));
result.put("serverPort", this.environment.getProperty("server.port", "8080"));
result.put("managementPort",
this.environment.getProperty("management.server.port", "(same as server.port)"));
result.put("managementBasePath",
this.environment.getProperty("management.endpoints.web.base-path", "/actuator"));
result.put("exposureInclude",
this.environment.getProperty("management.endpoints.web.exposure.include", "health"));
result.put("exposureExclude",
this.environment.getProperty("management.endpoints.web.exposure.exclude", "(none)"));
result.put("healthShowDetails",
this.environment.getProperty("management.endpoint.health.show-details", "never"));
Map<String, List<String>> exposed = new TreeMap<>();
for (ExposableWebEndpoint endpoint : this.webEndpoints.getEndpoints()) {
List<String> ops = new ArrayList<>();
for (WebOperation operation : endpoint.getOperations()) {
var predicate = operation.getRequestPredicate();
ops.add(predicate.getHttpMethod() + " " + predicate.getPath());
}
ops.sort(String::compareTo);
exposed.put(endpoint.getEndpointId().toString(), ops);
}
result.put("exposedWebEndpointCount", exposed.size());
result.put("exposedWebEndpoints", exposed);
List<String> contributors = new ArrayList<>();
collect("", this.healthRegistry, contributors);
contributors.sort(String::compareTo);
result.put("healthContributors", contributors);
return result;
}
private void collect(String prefix, HealthContributors contributors, List<String> into) {
for (HealthContributors.Entry entry : contributors) {
String name = prefix + entry.name();
HealthContributor contributor = entry.contributor();
if (contributor instanceof HealthContributors nested) {
collect(name + "/", nested, into);
}
else {
into.add(name + " (" + contributor.getClass().getSimpleName() + ")");
}
}
}
}

View File

@@ -0,0 +1,22 @@
package com.ankurm.actuator.web;
import org.springframework.jdbc.core.JdbcTemplate;
import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.RestController;
/** A minimal business endpoint, so the application is not only Actuator. */
@RestController
public class OrderController {
private final JdbcTemplate jdbc;
public OrderController(JdbcTemplate jdbc) {
this.jdbc = jdbc;
}
@GetMapping("/orders/count")
public String count() {
Integer n = this.jdbc.queryForObject("SELECT COUNT(*) FROM orders", Integer.class);
return "orders=" + n;
}
}

View File

@@ -0,0 +1,47 @@
package com.ankurm.actuator.web;
import com.ankurm.actuator.health.UpstreamState;
import org.springframework.http.ResponseEntity;
import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.PostMapping;
import org.springframework.web.bind.annotation.RequestMapping;
import org.springframework.web.bind.annotation.RequestParam;
import org.springframework.web.bind.annotation.RestController;
/**
* Stands in for the third-party API that {@code ExternalApiHealthIndicator} probes.
*
* <p>Keeping the "external" dependency inside the same JVM makes every failure in this
* repository reproducible offline and deterministic. A real deployment would point the
* indicator at a real host; nothing else about the indicator changes.
*
* <p>See docs/05-custom-health-indicators.md.
*/
@RestController
@RequestMapping("/stub/upstream")
public class StubUpstreamController {
private final UpstreamState state;
public StubUpstreamController(UpstreamState state) {
this.state = state;
}
@GetMapping("/ping")
public ResponseEntity<String> ping() throws InterruptedException {
switch (this.state.get()) {
case DOWN -> {
return ResponseEntity.status(503).body("upstream unavailable");
}
case SLOW -> Thread.sleep(30_000L);
case UP -> { }
}
return ResponseEntity.ok("pong");
}
@PostMapping("/mode")
public String mode(@RequestParam("value") String value) {
this.state.set(UpstreamState.Mode.valueOf(value.toUpperCase()));
return "upstream mode = " + this.state.get();
}
}

View File

@@ -0,0 +1,9 @@
management:
endpoints:
web:
exposure:
include: health,info,diag
endpoint:
health:
show-details: always
show-components: always

View File

@@ -0,0 +1,7 @@
# The configuration that appears in most tutorials, and the reason Actuator has a reputation
# for leaking. Exposes every endpoint the classpath provides on the web, including heapdump.
management:
endpoints:
web:
exposure:
include: "*"

View File

@@ -0,0 +1,38 @@
# Health groups: the fix for "a third-party outage restarted every pod we own".
#
# liveness -> is this JVM broken beyond recovery? Restarting is the only cure.
# readiness -> should this instance receive traffic right now?
#
# The external API and Kafka belong in readiness. They must never appear in liveness: a
# dependency being down is not a reason for the orchestrator to kill your process, and if it
# does, every instance restarts at once and you have turned a partial outage into a total one.
management:
endpoints:
web:
exposure:
include: health,info,diag
endpoint:
health:
show-details: always
group:
liveness:
include: livenessState,diskSpace
show-details: always
readiness:
include: readinessState,ordersDatabase,externalApi
show-details: always
# Kafka being down should degrade, not black-hole, this instance. OUT_OF_SERVICE
# still maps to 503 by default; the additional-path below is what the probe hits.
startup:
include: ordersDatabase
show-details: always
probes:
enabled: true
# Kafka is left out of the readiness group above ONLY so that this demonstration has one
# moving part. No broker runs in this repository, so including it would pin readiness to 503
# and hide the effect of the external API flipping. In a real service Kafka belongs in
# readiness alongside the database.
demo:
kafka:
bootstrap: localhost:9092

View File

@@ -0,0 +1,17 @@
# The textbook AdminClient health check: bootstrap servers and nothing else.
#
# demo.kafka.timeout-ms is set far above Kafka's own default on purpose. It is NOT what bounds
# the call here - the point of this profile is to let Kafka's default.api.timeout.ms (60000) be
# the binding constraint and to measure it. See docs/06-health-indicator-failure-modes.md.
demo:
kafka:
tuned: false
timeout-ms: 180000
management:
endpoints:
web:
exposure:
include: health,info,diag
endpoint:
health:
show-details: always

View File

@@ -0,0 +1,16 @@
# Actuator on its own port, on its own path.
#
# The point is not tidiness. It is that port 9001 can be bound to the pod network and left out
# of the ingress/load-balancer configuration entirely, so /actuator is unreachable from the
# internet by routing rather than by an authorisation rule you have to keep correct.
management:
server:
port: 9001
# Bind to loopback only. In Kubernetes you would leave this unset and simply not list 9001
# as a Service port; here it demonstrates that the address is separately controllable.
address: 127.0.0.1
endpoints:
web:
base-path: /manage
exposure:
include: "*"

View File

@@ -0,0 +1,6 @@
# Points the Kafka indicator at nothing at all, with a very short budget, so the DOWN path is
# fast. Used by the timeout demonstration.
demo:
kafka:
bootstrap: 10.255.255.1:9092
timeout-ms: 400

View File

@@ -0,0 +1,9 @@
# Pairs with OpenActuatorConfig. Deliberately bad. Do not copy.
management:
endpoints:
web:
exposure:
include: "*"
endpoint:
health:
show-details: always

View File

@@ -0,0 +1,13 @@
# Pairs with SecuredActuatorConfig.
management:
endpoints:
web:
exposure:
include: "*"
endpoint:
health:
# Anonymous callers get {"status":"UP"}. Authenticated ACTUATOR callers get the
# per-contributor breakdown. This is the setting that keeps a health endpoint useful to
# your operators without telling an attacker which of your dependencies is wobbling.
show-details: when-authorized
roles: ACTUATOR

View File

@@ -0,0 +1,65 @@
# Base configuration. Everything Actuator-related here is either a Spring Boot default written
# out explicitly (so you can see it) or a demo knob. Profile files layer on top.
spring:
application:
name: actuator-production
datasource:
url: jdbc:h2:mem:orders;DB_CLOSE_DELAY=-1
username: sa
password: "not-a-real-password-but-watch-what-/actuator/env-does-with-it"
jpa:
hibernate:
ddl-auto: none
sql:
init:
mode: always
security:
user:
name: ops
password: ops-password
roles: ACTUATOR
server:
port: 8080
demo:
upstream:
url: http://localhost:8080/stub/upstream/ping
timeout-ms: 750
# The stub is behind this application's own security. A real upstream would be behind
# someone else's. Either way the indicator must present credentials, or it reports DOWN
# for the wrong reason - see the note in ExternalApiHealthIndicator.
username: ops
password: ops-password
kafka:
bootstrap: localhost:9092
timeout-ms: 1500
# A custom property whose name does NOT match Spring Boot's sanitisation patterns.
# /actuator/env treats it exactly like any other value - see docs/03-endpoint-catalogue.md.
acme:
partner:
credential: "S3CRET-partner-credential"
management:
info:
env:
enabled: true
endpoint:
health:
# Boot's default. Spelled out so the contrast with the 'details' profile is visible.
show-details: never
info:
app:
name: actuator-production
purpose: companion repository for the ankurm.com Actuator article
logging:
level:
# The AdminClient's background thread is chatty when the broker is unreachable.
# Quietened here so docs/output/ transcripts stay readable; see the note in
# KafkaHealthIndicator about metadata.recovery.strategy.
org.apache.kafka.clients.admin.internals.AdminMetadataManager: WARN
org.apache.kafka.clients.NetworkClient: ERROR

View File

@@ -0,0 +1,4 @@
DELETE FROM orders;
INSERT INTO orders (id, customer, total_cents) VALUES (1, 'acme', 1999);
INSERT INTO orders (id, customer, total_cents) VALUES (2, 'globex', 24500);
INSERT INTO orders (id, customer, total_cents) VALUES (3, 'initech', 750);

View File

@@ -0,0 +1,5 @@
CREATE TABLE IF NOT EXISTS orders (
id BIGINT PRIMARY KEY,
customer VARCHAR(64) NOT NULL,
total_cents BIGINT NOT NULL
);

View File

@@ -0,0 +1,74 @@
package com.ankurm.actuator;
import org.junit.jupiter.api.Test;
import org.springframework.beans.factory.annotation.Autowired;
import org.springframework.boot.test.context.SpringBootTest;
import org.springframework.http.MediaType;
import org.springframework.test.context.ActiveProfiles;
import org.springframework.test.web.servlet.MockMvc;
import org.springframework.test.web.servlet.setup.MockMvcBuilders;
import org.springframework.web.context.WebApplicationContext;
import static org.springframework.security.test.web.servlet.setup.SecurityMockMvcConfigurers.springSecurity;
import static org.springframework.test.web.servlet.request.MockMvcRequestBuilders.get;
import static org.springframework.test.web.servlet.request.MockMvcRequestBuilders.post;
import static org.springframework.test.web.servlet.result.MockMvcResultMatchers.status;
/**
* Contract tests. These pin the surprising behaviour, not the happy path &mdash; if a future
* Spring Boot upgrade changes any of it, this suite is where you find out.
*
* <p>See docs/09-testing-actuator.md.
*/
@SpringBootTest
@ActiveProfiles({ "exposeall", "secured" })
class ActuatorExposureTests {
@Autowired
private WebApplicationContext context;
private MockMvc mvc() {
return MockMvcBuilders.webAppContextSetup(this.context).apply(springSecurity()).build();
}
@Test
void heapdumpIsNotExposedEvenWithWildcardExposure() throws Exception {
// management.endpoint.heapdump.access defaults to 'none'. Wildcard exposure does not
// override access. This is the assertion that fails loudest if someone "fixes" it.
mvc().perform(get("/actuator/heapdump")).andExpect(status().isNotFound());
}
@Test
void shutdownIsNotExposedEvenWithWildcardExposure() throws Exception {
mvc().perform(post("/actuator/shutdown")).andExpect(status().isNotFound());
}
@Test
void healthIsAnonymousButEnvIsNot() throws Exception {
mvc().perform(get("/actuator/health")).andExpect(status().isServiceUnavailable());
mvc().perform(get("/actuator/env")).andExpect(status().isUnauthorized());
}
@Test
void envMasksEveryValueRegardlessOfKeyName() throws Exception {
// Not just keys that look like secrets. show-values defaults to 'never'.
mvc().perform(get("/actuator/env/acme.partner.credential")
.with(org.springframework.security.test.web.servlet.request.SecurityMockMvcRequestPostProcessors
.user("ops").roles("ACTUATOR"))
.accept(MediaType.APPLICATION_JSON))
.andExpect(status().isOk())
.andExpect(org.springframework.test.web.servlet.result.MockMvcResultMatchers
.jsonPath("$.property.value").value("******"));
}
@Test
void anonymousHealthCarriesNoComponentBreakdown() throws Exception {
// show-details: when-authorized. An anonymous prober must not learn which dependency
// is failing.
mvc().perform(get("/actuator/health"))
.andExpect(status().isServiceUnavailable())
.andExpect(org.springframework.test.web.servlet.result.MockMvcResultMatchers
.jsonPath("$.components").doesNotExist());
}
}

View File

@@ -0,0 +1,81 @@
package com.ankurm.actuator;
import com.ankurm.actuator.health.UpstreamState;
import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.Test;
import org.springframework.beans.factory.annotation.Autowired;
import org.springframework.boot.health.contributor.Status;
import org.springframework.boot.health.registry.HealthContributorRegistry;
import org.springframework.boot.health.contributor.HealthIndicator;
import org.springframework.boot.test.context.SpringBootTest;
import org.springframework.test.context.ActiveProfiles;
import static org.assertj.core.api.Assertions.assertThat;
/**
* The contract that matters operationally: an upstream outage must move readiness and must NOT
* move liveness.
*
* <p>See docs/07-groups-and-probes.md.
*/
// DEFINED_PORT, not the default MOCK environment. ExternalApiHealthIndicator makes a real
// HTTP call to this application's own stub controller, so a servlet container has to be
// listening on the port the indicator was configured with. With the mock environment the
// indicator reports DOWN with "Connection refused" and the test proves nothing.
@SpringBootTest(webEnvironment = SpringBootTest.WebEnvironment.DEFINED_PORT)
@ActiveProfiles("groups")
class HealthGroupTests {
@Autowired
private HealthContributorRegistry registry;
@Autowired
private UpstreamState upstream;
@AfterEach
void reset() {
this.upstream.set(UpstreamState.Mode.UP);
}
private Status statusOf(String name) {
var contributor = this.registry.getContributor(name);
assertThat(contributor).as("contributor '%s' is registered", name).isNotNull();
return ((HealthIndicator) contributor).health().getStatus();
}
@Test
void everyCustomIndicatorIsRegisteredUnderTheExpectedName() {
// The bean name, minus a trailing "HealthIndicator", is the key in the JSON response.
// Rename the bean and you silently break every dashboard that reads it.
assertThat(statusOf("ordersDatabase")).isNotNull();
assertThat(statusOf("externalApi")).isNotNull();
assertThat(statusOf("kafka")).isNotNull();
}
@Test
void upstreamOutageMovesTheExternalApiIndicatorButNotLiveness() {
assertThat(statusOf("externalApi")).isEqualTo(Status.UP);
assertThat(statusOf("livenessState")).isEqualTo(Status.UP);
this.upstream.set(UpstreamState.Mode.DOWN);
assertThat(statusOf("externalApi")).isEqualTo(Status.DOWN);
assertThat(statusOf("livenessState"))
.as("a third-party outage must never make this process look unrecoverable")
.isEqualTo(Status.UP);
}
@Test
void theExternalApiIndicatorRespectsItsTimeoutBudget() {
this.upstream.set(UpstreamState.Mode.SLOW);
long start = System.nanoTime();
Status status = statusOf("externalApi");
long elapsedMs = (System.nanoTime() - start) / 1_000_000;
assertThat(status).isEqualTo(Status.DOWN);
// The stub sleeps 30s. demo.upstream.timeout-ms is 750. If this assertion ever fails,
// someone removed the read timeout and the health endpoint can now block a worker.
assertThat(elapsedMs).isLessThan(5_000L);
}
}