Files
asmhatreandClaude Sonnet 5 e478eafda3 Add db-migrations-expand-contract: zero-downtime schema migrations proven with a real 4-deploy rolling run
Companion code for Zero-Downtime Database Migrations: Expand-Contract in Practice
with Spring Boot: a full expand/migrate-writes/migrate-reads/contract sequence run
as an actual rolling deploy across two live replicas, with a load generator sending
continuous HTTP traffic through all four deploys (99.98% success, every residual
error traced to a root cause rather than left unexplained). Findings include a real
NOT NULL constraint trap in the expand migration, a backfill-window bug in the read
switch, H2's AUTO_SERVER=TRUE single-point-of-failure behavior under a rolling
restart, the drain-before-SIGTERM fix needed to close a health-check gap during
graceful shutdown, and H2 silently discarding a concurrently committed INSERT during
an ALTER TABLE ADD/DROP COLUMN rebuild - confirmed, by primary source, to be an
H2-specific behavior rather than a property of the technique itself.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_019Fb7vW8vLyLKngBc4R3huA
2026-09-16 19:22:48 +00:00

3.0 KiB

7. The backfill window bug

← 6. The NOT NULL trap · Next: 8. The rolling window proof →

Deploy 1's migration backfills email_address for every row that exists at the moment it runs. It cannot see rows written after that — and a rolling deploy is not instantaneous, so there's a real window, between "Deploy 1's migration finished" and "every replica in the fleet is confirmed Stage 2 or later", during which a still-live Stage 1 instance keeps inserting rows the old way: email only, email_address untouched, NULL.

That window isn't a hypothetical edge case — it's guaranteed to happen for however long Deploy 2's rollout takes, on every real fleet bigger than one instance.

BackfillWindowBugTest writes exactly such a row, then reads it back two ways:

-- the row a lingering Stage 1 instance just wrote --
[NAME=Katherine Johnson, [email protected], EMAIL_ADDRESS=null]

-- a NAIVE Stage 3 read (email_address alone) - the bug --
naive Stage 3 email column value: null

-- the SHIPPED Stage 3 read (CustomerService, COALESCE) - the fix --
CustomerService (stage 3) result: Optional[Customer[id=1, name=Katherine Johnson, [email protected]]]

Full transcript: docs/output/07-backfill-window-bug.txt.

The naive read — SELECT email_address FROM customers WHERE id = ? — is exactly what you'd write if you thought of the read switch as "just point at the new column now that it's backfilled". It's wrong specifically for rows born during Deploy 2's own rollout, because those rows are younger than Deploy 1's one-time backfill and were written by an instance that (correctly, for its stage) never touched email_address.

The fix is the COALESCE(email_address, email) already shown in chapter 5. It costs nothing for a fully dual-written row — COALESCE returns the first non-null argument, and both columns already agree — and it's the only thing that returns a correct answer for a row that hasn't caught up yet.

Going deeper

  • Stage 2 never needs this fallback: it still reads the original email column, which every stage keeps populated the entire time. Only Stage 3, the first stage to prefer the new column, needs the fallback — see the read switch in chapter 5.
  • This is the database-level version of a general rolling-deploy rule: any code path that reads data written by a different stage has to tolerate that stage's write shape, not just the shape your own stage would have produced.

← 6. The NOT NULL trap · Next: 8. The rolling window proof →