Companion code for Zero-Downtime Database Migrations: Expand-Contract in Practice with Spring Boot: a full expand/migrate-writes/migrate-reads/contract sequence run as an actual rolling deploy across two live replicas, with a load generator sending continuous HTTP traffic through all four deploys (99.98% success, every residual error traced to a root cause rather than left unexplained). Findings include a real NOT NULL constraint trap in the expand migration, a backfill-window bug in the read switch, H2's AUTO_SERVER=TRUE single-point-of-failure behavior under a rolling restart, the drain-before-SIGTERM fix needed to close a health-check gap during graceful shutdown, and H2 silently discarding a concurrently committed INSERT during an ALTER TABLE ADD/DROP COLUMN rebuild - confirmed, by primary source, to be an H2-specific behavior rather than a property of the technique itself. Co-Authored-By: Claude Sonnet 5 <[email protected]> Claude-Session: https://claude.ai/code/session_019Fb7vW8vLyLKngBc4R3huA
5.9 KiB
11. The load generator
← 10. What happens if you drop too soon · Next: 12. The AUTO_SERVER trap →
Every other chapter so far proves a property of the technique with a JUnit test
against a shared JdbcClient — real code, real SQL, no mocks, but also no real HTTP,
no real process restarts, no real timing pressure. scripts/run-all.sh is where all
three of those show up: two real Spring Boot processes, a real standalone database, a
real rolling restart between each stage, and
LoadGenerator
sending continuous HTTP traffic through the whole sequence.
What it does
Eight worker threads, each running the same loop: pick a healthy backend, then create a customer (50% of the time), read one it already knows about (30%), or update one (20%). Every read and update verifies the response body matches what the load generator itself expects — not just the HTTP status code:
JsonNode node = MAPPER.readTree(resp.body());
String returnedEmail = node.get("email").asString();
if (!target.email().equals(returnedEmail)) {
c.recordError("read-consistency-mismatch");
return;
}
A 200 with the wrong email in the body would be a much worse bug than a 500, and a naive load test that only checks status codes would never catch it.
Health-checked traffic, not a raw hose
LoadGenerator polls /actuator/health on both ports every 300ms and only sends
traffic to backends it currently believes are up:
private static final int UNHEALTHY_THRESHOLD = 2;
A backend needs two consecutive failed checks before it's removed from rotation —
this exists because this whole sequence runs two JVMs on a shared, small sandbox, and
one replica's cold-start CPU burst can make its sibling miss a single health check
without actually being down. Requiring two consecutive failures is the same debounce
a real load balancer's health check threshold gives you, and skipping it turned
transient slowness into false no-healthy-backend errors in an earlier version of
this test.
That fix mattered enough to also show up in
start-instance.sh,
on the other side of the same problem — cutting each replica's own startup CPU cost
so it's less likely to starve its sibling in the first place:
nohup java -XX:TieredStopAtLevel=1 -XX:+UseSerialGC -Xms128m -Xmx256m \
-jar "$JAR" --server.port="$PORT" --app.stage="$STAGE" \
Per-thread state, not shared state
Each worker thread owns a private ArrayDeque of customer ids it has created — never
shared with the other seven threads:
workers.submit(() -> {
try {
workerLoop(new ArrayDeque<>());
An earlier version shared one pool across all eight threads, and produced
read-consistency-mismatch errors that had nothing to do with the server at all: two
threads racing to update the same shared id could leave the pool holding a stale
expected value, so a perfectly correct server response looked like a bug. Giving each
thread exclusive ownership of the rows it creates removes that entire class of false
positive while still hammering both replicas concurrently.
The result
Load generator summary
=======================
Total requests: 30911
Successful: 30905
Errors: 6
By phase:
04a-stage4-soak ok=2130 error=0
04b-contract-migration ok=336 error=0
04a-deploy-stage4-rollout ok=4147 error=0
03-stage3-soak ok=2116 error=0
03-deploy-stage3-rollout ok=4645 error=0
00-baseline-soak ok=1893 error=0
05-final-soak ok=7118 error=4
- read-http-404 2
- update-http-404 2
02-stage2-soak ok=2068 error=0
01-expand-migration ok=1773 error=1
- read-http-404 1
02-deploy-stage2-rollout ok=4679 error=1
- update-http-404 1
Full transcript:
docs/output/12-load-generator-summary.txt,
and the phase-by-phase deploy log this run came from:
docs/output/11-live-deploy-sequence.txt.
30,905 of 30,911 requests succeeded — 99.98%, across a real rolling restart through
all four deploys. The six errors are all 404s, not 500s: a request for a customer
id that genuinely wasn't found, not a crash. Every one of them traces to the same
root cause, and it's the most interesting finding in this whole module — see
chapter 14.
Going deeper
- The graceful-shutdown-and-drain sequence that gets the rolling restart portion of this run down to zero client-visible errors on its own is chapter 13 — the six remaining errors above have a different cause entirely, isolated in chapter 14.
- Every deploy's schema state during this exact run, captured live:
docs/output/13-schema-diagnostics-timeline.txt, viaSchemaDiagnosticsController.
← 10. What happens if you drop too soon · Next: 12. The AUTO_SERVER trap →