1
0
Files
zgc-jdk25-benchmarks/docs/08-choosing.md
Ankur e189da79ba Generational ZGC vs G1 on JDK 25: benchmarks, internals and captured results
Complete companion suite for the ankurm.com guide. Everything was compiled and
executed on java 25.0.3+9-LTS-195 (AMD Ryzen 5 5600U, Windows 11); every file
under results/ is unedited program output.

Code
  latency/    open-loop latency harness that measures from intended arrival, so
              coordinated omission is accounted for rather than hidden. Reports
              response time and service time side by side; the gap reached 2910x
              on G1.
  jmh/        four microbenchmarks isolating one mechanism each: TLAB allocation,
              the ZGC load barrier (with a primitive-load control), the write
              barrier (with null / old-to-old / primitive controls), and promotion
              pressure.
  tuning/     provokes ZGC allocation stalls and reads its own JFR recording back,
              grouped by page class. jdk.ZAllocationStall is enabled without a
              threshold, because the 10 ms default hides most stalls.
  internals/  ZGC page classes vs G1 humongous, read from the live VM via
              HotSpotDiagnosticMXBean and jdk.ZPageAllocation events.
  analysis/   unified GC log parser that keeps stop-the-world pauses and
              concurrent phases in separate buckets.
  env/        environment capture; proves generational mode from JMX bean names.

Docs
  Eight chapters covering the JEP 439/474/490 timeline, colored pointers and both
  barriers, allocation stalls, a full flag reference, logging and JFR, benchmark
  methodology, corner cases, and a decision procedure.

Headline result (60 s, 20k req/s, 2 GB heap, ~585 MB live)
  p50               ZGC 0.006 ms   G1 0.005 ms
  p99.9             ZGC 1.437 ms   G1 95.169 ms
  total STW time    ZGC 1.503 ms   G1 1,139.624 ms
2026-07-31 08:47:16 +05:30

5.1 KiB
Raw Permalink Blame History

8. Choosing a collector on JDK 25

8.1 The short version

G1 is still the right default. It is the default for a reason: it is balanced, it needs no tuning to be adequate, and it wins the median. Switch to ZGC when you have a stated tail-latency requirement that G1 is measurably missing — not because ZGC is newer.

8.2 What the measurements actually say

From the 60-second run in this repository (Ryzen 5 5600U, Windows 11, JDK 25.0.3, 2 GB heap, ~585 MB live, 20,000 req/s, ~328 MB/s allocation rate):

Metric ZGC G1 Winner
p50 response time 0.006 ms 0.005 ms G1, barely
p90 0.009 ms 0.009 ms tie
p99 0.029 ms 8.042 ms ZGC, 277x
p99.9 1.437 ms 95.169 ms ZGC, 66x
p99.99 23.675 ms 126.949 ms ZGC, 5.4x
max 32.178 ms 133.018 ms ZGC
mean 0.016 ms 0.388 ms ZGC
total stop-the-world time 1.503 ms 1,139.6 ms ZGC, 758x
longest single pause 0.054 ms 104.449 ms ZGC
Full GCs / evacuation failures 0 / 0 4 / 16 ZGC

Note the shape: the two collectors are indistinguishable up to p90. Everything ZGC buys you is in the tail. If your SLO is a mean or a p95, this entire table is telling you to stay on G1 and spend the effort elsewhere.

Note also that ZGC is not magic: its own p99.99 is 23.7 ms. Those are allocation stalls (the GC log for that run contains 40 of them), not pauses. ZGC moved the latency source, it did not delete it.

8.3 A decision procedure

Work through it in order. Stop at the first "yes".

  1. Is GC actually your problem? Get a JFR recording and look at jdk.GCPhasePause and jdk.ZAllocationStall against your request latency. Most "GC problems" are lock contention, downstream calls, or connection-pool starvation. Changing collector will not fix those and will cost you a week.
  2. Is your heap small (< 4 GB) and your SLO a mean or p95? Stay on G1. ZGC's advantages need a tail requirement and a heap large enough that marking cost matters.
  3. Do you have a hard tail-latency SLO — p99 or p99.9 in single-digit or low-double-digit milliseconds — that G1 is missing? This is the case ZGC was built for. Go to step 5.
  4. Is your heap very large (> 32 GB) with a large live set? G1's pause time scales with live data in the collection set; ZGC's does not. ZGC becomes increasingly favourable as the heap grows. Go to step 5.
  5. Can you afford the footprint? ZGC needs headroom (§8.4) and loses compressed oops (07-corner-cases.md §7.1). Budget 3050% more memory than the equivalent G1 deployment, more if your heap is reference-dense. If you cannot, ZGC will stall and you will conclude — wrongly — that ZGC is bad.
  6. Is your workload throughput-bound with a hot pointer-chasing inner loop? ZGC's load barrier is a real, permanent tax on reference loads. Batch jobs, analytics, and anything that walks large object graphs should measure before switching. Consider Parallel GC instead — for a batch job with no latency requirement, Parallel still beats both on raw throughput.

8.4 Sizing ZGC

The single most common ZGC failure is a heap sized as if it were a G1 heap.

  • Start at 1.5x the -Xmx you used for G1, then measure.
  • Set SoftMaxHeapSize to roughly 75% of -Xmx so there is a reserve for allocation spikes.
  • Watch jdk.ZAllocationStall (threshold 0 ms), not pause time. If stalls are non-zero under normal load, you are under-provisioned.
  • If stalls are concentrated on Large pages, the fix is in your application's allocation sizes, not in the heap size.

8.5 When the answer is neither

  • Parallel GC (-XX:+UseParallelGC) — batch jobs, ETL, anything where wall-clock throughput is the only metric and a multi-second pause is irrelevant. Still the throughput champion.
  • Serial GC (-XX:+UseSerialGC) — small containers, single-core, short-lived processes. Lowest overhead, smallest footprint, fastest start-up.
  • Shenandoah (-XX:+UseShenandoahGC) — the other concurrent-compacting collector. Broadly similar goals to ZGC with a different barrier design (Brooks-style forwarding rather than colored pointers); it keeps compressed oops, which makes it interesting precisely where §7.1 hurts. Worth benchmarking as a third option if footprint is your constraint.

8.6 The honest summary

ZGC does not make garbage collection free. It converts a rare, large, correlated cost — a stop-the-world pause that hits every thread at once — into a constant, small, uncorrelated cost: barrier work on every reference access, plus a memory footprint premium.

For a request-serving system with a tail-latency SLO, that trade is close to unambiguously good, and JDK 25 is the first LTS where you get it without an opt-in flag or a non-generational fallback to worry about.

For a batch job, it is a straightforward loss.

Most systems are somewhere in between, which is why the only defensible answer is to run scripts/run-all.ps1 against a workload shaped like yours.


Back to the README.