1
0
Files
zgc-jdk25-benchmarks/docs/06-methodology.md
Ankur e189da79ba Generational ZGC vs G1 on JDK 25: benchmarks, internals and captured results
Complete companion suite for the ankurm.com guide. Everything was compiled and
executed on java 25.0.3+9-LTS-195 (AMD Ryzen 5 5600U, Windows 11); every file
under results/ is unedited program output.

Code
  latency/    open-loop latency harness that measures from intended arrival, so
              coordinated omission is accounted for rather than hidden. Reports
              response time and service time side by side; the gap reached 2910x
              on G1.
  jmh/        four microbenchmarks isolating one mechanism each: TLAB allocation,
              the ZGC load barrier (with a primitive-load control), the write
              barrier (with null / old-to-old / primitive controls), and promotion
              pressure.
  tuning/     provokes ZGC allocation stalls and reads its own JFR recording back,
              grouped by page class. jdk.ZAllocationStall is enabled without a
              threshold, because the 10 ms default hides most stalls.
  internals/  ZGC page classes vs G1 humongous, read from the live VM via
              HotSpotDiagnosticMXBean and jdk.ZPageAllocation events.
  analysis/   unified GC log parser that keeps stop-the-world pauses and
              concurrent phases in separate buckets.
  env/        environment capture; proves generational mode from JMX bean names.

Docs
  Eight chapters covering the JEP 439/474/490 timeline, colored pointers and both
  barriers, allocation stalls, a full flag reference, logging and JFR, benchmark
  methodology, corner cases, and a decision procedure.

Headline result (60 s, 20k req/s, 2 GB heap, ~585 MB live)
  p50               ZGC 0.006 ms   G1 0.005 ms
  p99.9             ZGC 1.437 ms   G1 95.169 ms
  total STW time    ZGC 1.503 ms   G1 1,139.624 ms
2026-07-31 08:47:16 +05:30

6.5 KiB

6. Methodology — how to not publish a meaningless GC benchmark

Most GC comparisons on the internet are wrong in one of a small number of specific, fixable ways. This chapter lists them, because a benchmark's credibility is entirely a function of which of these it avoided.

6.1 Coordinated omission

The mistake. A closed-loop harness — while (true) { doRequest(); } — stops issuing load exactly when the system is least able to serve it. If a 100 ms pause freezes the application, the load generator does not send requests during those 100 ms. It records one slow request per thread and moves on. The pause was real; your users experienced it as 100 ms of queueing across ~2000 requests; the benchmark reports it as a handful of samples.

The fix. Fix the arrival schedule up front and measure from intended arrival, not from when the thread actually got around to the request.

long intended = start + i * interArrivalNanos;
parkUntil(intended);
long actualStart = System.nanoTime();
handleRequest();
long done = System.nanoTime();

responseTime.record(done - intended);      // honest
serviceTime.record(done - actualStart);    // what a closed-loop harness reports

How much it matters. This repository records both and prints the ratio. From the 60-second run on the reference machine:

Collector p99.9 response time p99.9 service time ratio
ZGC 1.437 ms 0.143 ms 10x
G1 95.169 ms 0.033 ms 2910x

A closed-loop benchmark of this exact workload would have concluded that G1's p99.9 is 0.033 ms — better than ZGC's 0.143 ms — and published a chart showing G1 winning on tail latency. The honest numbers say G1's p99.9 is 66x worse.

Note that the distortion is not uniform: it is far larger for the collector with larger pauses. So coordinated omission does not merely add noise, it systematically favours the collector you are trying to evaluate against.

6.2 Benchmarking the sleep instead of the system

The first version of this harness used LockSupport.parkNanos(remaining) between requests. On Windows a park rounds up to roughly 1 ms, and the inter-arrival interval was 200 µs. Every park overshot by several intervals; the thread then blasted through the backlog and parked again. The harness reported a p50 of 0.590 ms for a request that takes 3 µs, and the shape of the distribution looked plausibly like a GC effect.

The fix is to park only when the deadline is far away (3 ms here, comfortably beyond the timer slack) and busy-spin the rest. It costs a core per worker thread, which is why the worker count defaults to a third of the logical CPUs — see §6.4.

If your latency harness's p50 is much larger than the work it does, suspect your sleep before you suspect the collector.

6.3 Comparing pause time against cycle time

Covered in detail in 05-logging-and-jfr.md. In one line: ZGC's GarbageCollectorMXBean "Cycles" beans and its Concurrent ... log lines are not freeze time, and summing them against G1's pause numbers produces a comparison that is off by two orders of magnitude in G1's favour.

6.4 Starving the collector of CPU

ZGC does concurrent work on its own threads. If the benchmark's load generator occupies every core — which a busy-spinning open-loop harness will happily do — those threads cannot run, the collector falls behind, and you measure a scheduling problem while believing you measured a GC.

This repository's harness therefore defaults to availableProcessors() / 3 worker threads and states so in its output. On the 12-thread reference machine that is 4 spinning workers, leaving 8 logical CPUs for the JVM's own threads.

6.5 Letting the heap resize during measurement

-Xms-Xmx means the heap grows and shrinks during the run, and you are partly measuring the resize policy. Every run in this repository uses -Xms2g -Xmx2g.

A side effect worth knowing: ZUncommit automatically becomes false when -Xms == -Xmx (verified in results/05-pages-zgc.txt). Fixing the heap therefore also removes uncommit behaviour from the comparison. That is the right call for a benchmark and the wrong call if uncommit is the thing you wanted to study.

6.6 Forgetting that JMH forks a fresh JVM

This one silently invalidates a large fraction of published GC microbenchmarks:

# WRONG -- the fork does not inherit the outer JVM's GC flags
java -XX:+UseZGC -jar benchmarks.jar

# RIGHT
java -jar benchmarks.jar -jvmArgs "-XX:+UseZGC -Xms2g -Xmx2g"

JMH starts a separate JVM for each trial. Flags on the outer java command configure the harness, not the forked JVM under measurement. Run it the wrong way twice and you get two runs of the default collector, a pleasingly small difference between them, and a blog post.

Verify with -XX:+PrintFlagsFinal inside the fork, or simply read JMH's own # VM options: banner — it prints the fork's arguments.

6.7 Measuring an empty heap

A collector with nothing to trace is fast. A benchmark that only allocates immediately-dead garbage measures the TLAB allocation path and nothing else — and every collector has essentially the same TLAB path.

The interesting variable is the live set: how much genuinely reachable data must be marked, and how often it is mutated. This repository's harness holds ~585 MB live in a 2 GB heap and rewrites a random slot on 20% of requests. PromotionBench parameterises the same idea directly through survivorDepth.

6.8 Dead-code elimination

If the JIT can prove your allocation is unused, it deletes it, and you benchmark an empty loop. Use JMH's Blackhole or return values, and — critically — check gc.alloc.rate.norm from -prof gc. It reports bytes allocated per operation. If that number is not what the source code implies, escape analysis removed something and the benchmark is measuring nothing.

6.9 One run is not a measurement

Everything in results/ is a single run on a single laptop, and is presented as an illustration of shape, not as a benchmark result you should plan capacity from. Ryzen 5 5600U, Windows 11, 15.3 GB RAM, JDK 25.0.3 — with browser tabs, an OS, and thermal limits. Server hardware, Linux, and a quiet machine will produce different absolute numbers.

What generalises is the relationship: G1 wins the median, ZGC wins the tail, and the gap widens as the live set grows. What does not generalise is any specific millisecond figure.


Next: 07-corner-cases.md.