1
0
Files
zgc-jdk25-benchmarks/docs/07-corner-cases.md
Ankur e189da79ba Generational ZGC vs G1 on JDK 25: benchmarks, internals and captured results
Complete companion suite for the ankurm.com guide. Everything was compiled and
executed on java 25.0.3+9-LTS-195 (AMD Ryzen 5 5600U, Windows 11); every file
under results/ is unedited program output.

Code
  latency/    open-loop latency harness that measures from intended arrival, so
              coordinated omission is accounted for rather than hidden. Reports
              response time and service time side by side; the gap reached 2910x
              on G1.
  jmh/        four microbenchmarks isolating one mechanism each: TLAB allocation,
              the ZGC load barrier (with a primitive-load control), the write
              barrier (with null / old-to-old / primitive controls), and promotion
              pressure.
  tuning/     provokes ZGC allocation stalls and reads its own JFR recording back,
              grouped by page class. jdk.ZAllocationStall is enabled without a
              threshold, because the 10 ms default hides most stalls.
  internals/  ZGC page classes vs G1 humongous, read from the live VM via
              HotSpotDiagnosticMXBean and jdk.ZPageAllocation events.
  analysis/   unified GC log parser that keeps stop-the-world pauses and
              concurrent phases in separate buckets.
  env/        environment capture; proves generational mode from JMX bean names.

Docs
  Eight chapters covering the JEP 439/474/490 timeline, colored pointers and both
  barriers, allocation stalls, a full flag reference, logging and JFR, benchmark
  methodology, corner cases, and a decision procedure.

Headline result (60 s, 20k req/s, 2 GB heap, ~585 MB live)
  p50               ZGC 0.006 ms   G1 0.005 ms
  p99.9             ZGC 1.437 ms   G1 95.169 ms
  total STW time    ZGC 1.503 ms   G1 1,139.624 ms
2026-07-31 08:47:16 +05:30

164 lines
8.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 7. Corner cases and things that bite
The blog post covers the headline concepts. This chapter is the long tail — the behaviours that do
not fit a narrative but do show up in production. Every claim here was verified on
`java 25.0.3+9-LTS-195`; raw dumps in [`results/`](../results).
## 7.1 ZGC turns off compressed oops. This is the biggest hidden cost.
On a heap under ~32 GB, HotSpot normally stores references as 32-bit offsets rather than 64-bit
pointers ("compressed oops"), roughly halving the space taken by reference fields. **ZGC cannot do
this**, because it needs the high bits of every pointer for the color. Verified:
```console
# ZGC, 2 GB heap
bool UseCompressedOops = false {product lp64_product} {ergonomic}
# G1, same 2 GB heap
bool UseCompressedOops = true {product lp64_product} {ergonomic}
```
Consequences, in rough order of how often they surprise people:
- **Every reference field costs 8 bytes instead of 4.** A `HashMap.Node` (hash, key, value, next) goes
from 32 bytes to 48. A reference-dense heap — many small objects, many pointers, think a large
in-memory graph, an ORM session cache, a JSON DOM — can grow **3040%** just by switching
collector, with no application change.
- **Object header layout differs**, so heap-size estimates from a G1 heap dump do not transfer.
- **`-Xmx` is not comparable across collectors.** "We gave both 8 GB" is not a fair test if one of
them needs 11 GB to hold the same object graph. This repository uses a byte-array-heavy live set
partly for this reason: `byte[]` payloads are dominated by data, not references, which keeps the
comparison honest. A pointer-heavy live set would have handed ZGC a footprint penalty that has
nothing to do with pause behaviour.
- `UseCompressedClassPointers` stays `true` on both, so class pointers are unaffected.
If you are evaluating ZGC and your heap is full of small objects, measure footprint before you
measure latency.
## 7.2 ZGC large pages are never relocated
ZGC's Large pages hold exactly one object each and are **freed, not compacted**. A heap with many
long-lived multi-megabyte objects therefore fragments in a way ZGC structurally cannot fix. G1 has
the mirror problem with humongous regions, but G1 has a Full GC that *can* compact them (at
enormous cost).
Verified placement on a 2 GB heap — note the last line, which is the one that catches people:
```console
ZGC pages actually allocated while creating the objects above
page type count total bytes
Large 2 25,165,824
Medium 1 33,554,432
```
Two Large pages for the 4 MB and 16 MB allocations. **The 4 MB one is the surprise:** the documented
medium-object limit is 4 MB, but `new byte[4 * 1024 * 1024]` is 4 MB *plus a 16-byte array header*,
which is over the limit. It gets an 8 MB Large page to itself — a 2x footprint amplification caused
by sixteen bytes. If you pool buffers, pool them at `4 MB - 64` rather than exactly 4 MB.
## 7.3 G1's humongous threshold is lower than people think
At a 2 GB heap, `G1HeapRegionSize` is ergonomically 1 MB, so **anything over 512 KB is humongous**:
allocated straight into the old generation, spanning whole regions, never eligible for the cheap
young-collection path.
```console
G1HeapRegionSize : 1,048,576 bytes (1 MB)
=> G1 humongous threshold: objects larger than 524,288 bytes (512 KB) go straight to old gen
```
A service that reads 1 MB request bodies into a single `byte[]` is producing humongous garbage on
every request. `-XX:G1HeapRegionSize=8m` moves the threshold to 4 MB and can be transformative.
This is one of the few G1 flags that regularly delivers a large win, and it is invisible unless you
go looking.
## 7.4 Virtual thread stacks live on the heap
A mounted virtual thread runs on a carrier's stack, but an **unmounted** one has its stack copied
into the heap as `StackChunk` objects. So a service with 200,000 parked virtual threads is holding
200,000 medium-sized, medium-lived objects that the collector must trace.
This interacts badly with a naive reading of "virtual threads are cheap": they are cheap in
*scheduler* terms and not free in *GC* terms. Applications that move from a 200-thread pool to
200,000 virtual threads often see their live set and promotion rate change shape entirely, which is
exactly the regime where the ZGC-vs-G1 answer flips. If you have made that migration, re-run your GC
choice; the decision you made before it is stale.
## 7.5 `System.gc()` behaves differently
`ExplicitGCInvokesConcurrent` is `false` by default on **both** collectors here. Under G1 that means
an explicit `System.gc()` triggers a **Full GC** — a stop-the-world compaction. Under ZGC it
triggers a major cycle, most of which is still concurrent, so the blast radius is much smaller.
Libraries that call `System.gc()` (some direct-`ByteBuffer` cleanup paths, some test harnesses, some
profilers) are therefore far more damaging on G1. `-XX:+ExplicitGCInvokesConcurrent` is a cheap
mitigation on G1; `-XX:+DisableExplicitGC` is the blunt one.
Note that `StoreBarrierBench` in this repository calls `System.gc()` twice in `@Setup` — deliberately,
to force the array into the old generation before measurement, and outside the measured window.
## 7.6 Reference processing still stops the world briefly
Weak, soft, phantom references and finalizers are processed during specific phases. On ZGC this
happens under `Concurrent Process Non-Strong`, visible in the log:
```
Concurrent Process Non-Strong n=8 total= 5.635 ms max= 0.984 ms
```
Concurrent, therefore not freeze time — but a heap with millions of weak references still costs
real CPU and can extend cycle duration enough to trigger allocation stalls indirectly. `WeakHashMap`
at scale, and caches built on `SoftReference`, are the usual suspects. Soft references in particular
interact with `SoftRefLRUPolicyMSPerMB` and are cleared based on *free heap*, which behaves very
differently when `SoftMaxHeapSize` is set.
## 7.7 `ZUncommit` disables itself on a fixed heap
```console
# -Xms2g -Xmx2g
ZUncommit = false
```
Logical — there is nothing to give back — but it surprises teams who set a fixed heap for
predictability *and* expect memory returned to the container. If you want uncommit, you must give
`-Xms` and `-Xmx` different values, and tune `ZUncommitDelay` (default 300 s) down if you want it to
happen promptly.
## 7.8 GC thread counts differ by default
```console
ZGC : ParallelGCThreads = 8, ConcGCThreads = 3
G1 : ParallelGCThreads = 10, ConcGCThreads = 3
```
On the same 12-thread machine. If you pin GC threads for consistency across a fleet, you are
overriding two different ergonomic decisions with one number, and one of the collectors will be
mis-sized. Prefer leaving them ergonomic unless you are running in a CPU-limited container, where
you should set them explicitly because the JVM's view of `availableProcessors()` may not match the
cgroup quota.
## 7.9 Container limits
Both collectors respect cgroup limits, but the failure modes differ:
- ZGC's higher footprint (see §7.1) plus its need for headroom means a container sized for G1 will
push ZGC into allocation stalls. Budget more memory, not the same memory.
- `-XX:MaxRAMPercentage` is the right knob in containers; `-Xmx` hard-codes a number that will be
wrong after the next resize.
- ZGC reserves a large *virtual* address range. Tools that alarm on virtual size rather than RSS
will report frightening numbers. Alarm on RSS.
## 7.10 Things that are the same and often assumed not to be
- **String deduplication** works on ZGC (since JDK 18), not just G1. `-XX:+UseStringDeduplication`
is off by default on both.
- **Class unloading** happens on both, during major/concurrent cycles.
- **TLABs** work identically; the fast allocation path is a pointer bump on both.
- **Heap dumps** (`jcmd GC.heap_dump`) work the same and are stop-the-world on both.
- **`-XX:+AlwaysPreTouch`** is `false` by default on both, and is worth turning on for both if you
care about first-minute latency.
---
Next: [08-choosing.md](08-choosing.md).