vt-pinning: diagnosing virtual thread pinning companion code

Reproduces the pinning JEP 491 did not fix (native JNI frames) with a real
native library, alongside the monitor case it did fix. Covers the
jdk.VirtualThreadPinned JFR event, reading pinned vs unmounted virtual
threads from a jcmd Thread.dump_to_file, the measured throughput cost on a
capped carrier pool, and a bounded-executor pattern to contain the blast
radius.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FhzLY5p6okFva3qsnsRyvM
This commit is contained in:
2026-09-30 06:35:35 +00:00
co-authored by Claude Sonnet 5
parent ff183da163
commit 0819993c51
18 changed files with 725 additions and 0 deletions
+90
View File
@@ -0,0 +1,90 @@
# vt-pinning
Companion code for the ankurm.com post *"Diagnosing Virtual Thread Pinning in Production: JFR
Events, jcmd, and Real Fixes."* Fourth module in `java-core-examples`, the Java-core /
concurrency series.
[JEP 491](https://openjdk.org/jeps/491) (GA in JDK 24) fixed the pinning caused by `synchronized`.
It did not touch the pinning caused by native code - a virtual thread executing a JNI native
method (or an FFM downcall) whose native frame calls back into blocking Java code still pins its
carrier, on every JDK version, because a continuation cannot be frozen across a native stack
frame. This module reproduces that with a real, tiny JNI library (not asserted from documentation)
and works through the diagnostics that still catch it once the old `-Djdk.tracePinnedThreads` flag
stops reporting anything at all.
## Versions this was built and tested against
| Component | Version | Notes |
|---|---|---|
| JDK (primary) | 25.0.4.1+1 (Temurin, LTS) | Has JEP 491; native pinning still reproduces. |
| JDK (comparison) | 21.0.10 (OpenJDK) | Pre-JEP-491, for the before/after. |
| gcc | system default | Builds the JNI native library; not committed, built by `scripts/build-native.sh`. |
| Maven | 3.9.11 | Compiles the Java classes only - no JMH in this module. |
| Hardware | 2 vCPU x86-64 VM | Same sandbox as the rest of this series. |
## Quickstart
```bash
export JDK25_HOME=/path/to/jdk-25
export JDK21_HOME=/path/to/jdk-21
export JAVA_HOME=$JDK21_HOME # any JDK's jni.h works for the native build
./scripts/build-native.sh
$JDK25_HOME/bin/javac --release 25 -d target/classes $(find src/main/java -name "*.java")
$JDK25_HOME/bin/java --enable-native-access=ALL-UNNAMED -Djava.library.path=target/native \
-cp target/classes com.ankurm.vtpinning.NativePinningDemo
```
`scripts/run-all.sh` regenerates every file in `output/` (needs both `JDK25_HOME` and
`JDK21_HOME` set, plus `gcc` on `PATH`).
## What's in here
| File | What it shows |
|---|---|
| `src/main/c/vtpinning.c` | One JNI native function, shared by every demo class: calls back into a static Java method on the same native stack frame, so the callback's blocking happens with native frames still present. |
| `.../MonitorPinningDemo.java` | `synchronized` + `Thread.sleep` on a virtual thread - the case JEP 491 fixed. |
| `.../NativePinningDemo.java` | The case JEP 491 did not fix: a JNI native method whose callback blocks. |
| `.../ThreadDumpPinnedDemo.java` | One pinned and one unpinned virtual thread running concurrently, for a `jcmd Thread.dump_to_file` capture. |
| `.../PinningThroughputDemo.java` | The wall-time cost of native pinning vs. plain blocking, on a capped carrier pool. |
| `.../BoundedNativeCallDemo.java` | The fix: routing a pinning native call through a small dedicated executor instead of calling it directly from the shared carrier pool. |
| `output/01` | `MonitorPinningDemo` on JDK 21 vs JDK 25, with `-Djdk.tracePinnedThreads=full`. |
| `output/02` | `NativePinningDemo` on JDK 21 vs JDK 25, same flag - still fires on 21, silent on 25. |
| `output/03` | The `jdk.VirtualThreadPinned` JFR event, captured and printed with `jfr print`, for the same native pin on JDK 25. |
| `output/04` | A `jcmd Thread.dump_to_file -format=json` snapshot taken mid-pin, pinned and unpinned virtual thread entries side by side. |
| `output/05` | `PinningThroughputDemo`: 4 virtual threads / 2 carriers, native vs plain. |
| `output/06` | `BoundedNativeCallDemo`: direct vs bounded-executor routing, effect on unrelated work. |
## Reading the numbers honestly (2-vCPU sandbox)
**The old diagnostic flag doesn't just go quiet for the case JEP 491 fixed - it's silent for the
case JEP 491 left alone too** (`output/02`). `-Djdk.tracePinnedThreads=full` reports
`reason:NATIVE` on JDK 21 for the exact same native call that still measurably pins on JDK 25
(confirmed independently via the JFR event in `output/03`, and via the throughput cost in
`output/05`). A team that kept the flag in their JVM args through an upgrade to JDK 24+ gets no
signal from it for either kind of pinning going forward - not just the kind that stopped
mattering.
**A `jcmd` thread dump distinguishes pinned from unmounted with one field** (`output/04`): a
pinned virtual thread's JSON entry carries a `"carrier"` field naming the platform thread's `tid`,
and its stack includes `VirtualThread.parkOnCarrierThread`. An unmounted (not pinned) blocked
virtual thread has neither - no `carrier` field at all, and a plain `parkNanos` in its stack. This
needs no JFR recording running and works against a single dump taken after the fact.
**Native pinning genuinely serializes work on a capped carrier pool** (`output/05`): 4 virtual
threads each blocking ~1000ms on 2 carriers finish in ~2015ms through the native path (two
carrier-bound batches) versus ~1022ms through plain `Thread.sleep` (all four unmount and share the
2 carriers freely). This isn't a benchmark artifact - it's the same mechanism JEP 491 fixed for
monitors, just for a case it didn't touch.
**Routing the pinning call through a small dedicated executor protects everything else**
(`output/06`): with 2 carriers, 2 concurrent native-pinning calls, and 6 unrelated
100ms-sleep virtual threads competing for the same pool, calling the native code directly starves
the unrelated work until ~924ms - both carriers are pinned the whole time. Routing the same native
calls through a 2-thread dedicated executor (via `Future.get()`, which parks normally with no
native frame on *that* thread's own stack) lets the unrelated work finish in ~126ms while the
native calls run to completion in the background. The pinning cost doesn't go away, but it stops
being everyone else's problem.
## License
MIT - see the [repo-wide LICENSE](../LICENSE).