Add virtual-threads-benchmark-webflux: the WebFlux leg of the three-way benchmark

Fixes found during self-correction before publishing:
- /stream used Flux.interval(), which ticks on its own wall-clock schedule
  independent of downstream demand and threw OverflowException under a slow
  subscriber; switched to Flux.range(), which has no independent production
  schedule and can never outrun demand.
- Single-trial HTTP load tests on this shared sandbox swung by more than 50%
  run to run (795ms-1247ms observed on the identical /io endpoint back to
  back) -- large enough to flip which threading model looked faster. Fixed
  by taking the median of 5 independent trials for the I/O-bound benchmark
  and the median of 3 for the event-loop-starvation benchmark, rather than
  reporting a single noisy run as if it were precise.
- The event-loop-starvation test's first cut used only 8 concurrent /cpu
  requests as background load, which drained through the 4 event-loop
  threads well inside the /io measurement window and produced an
  inconsistent, sometimes-inverted result across runs; raising to 60 fixed
  the under-loading problem but still flaked once during verification
  (372ms vs 374ms p99, a real tie). Final fix: 150 concurrent requests plus
  the median-of-3 trials above.

Also adds StreamBackpressureTest, a StepVerifier proof that the /stream
endpoint never emits ahead of its subscriber's outstanding requests, and
updates the module's docs to report the de-noised numbers with an explicit
methodology note on how they compare to the single-trial platform/virtual-
thread numbers reused from a different post.
This commit is contained in:
Claude
2026-09-19 09:34:18 +00:00
parent f506b01389
commit 09631dcaab
19 changed files with 975 additions and 0 deletions
@@ -0,0 +1,187 @@
# 1. WebFlux benchmark methodology and results
[README](../README.md) | Companion module: [`../virtual-threads-benchmark`](../virtual-threads-benchmark/README.md) (platform threads vs virtual threads)
Source: [`ReactiveDemoController.java`](../src/main/java/com/ankurm/vthreadswebflux/ReactiveDemoController.java), [`CpuWork.java`](../src/main/java/com/ankurm/vthreadswebflux/CpuWork.java).
Test: [`WebfluxLoadBenchmarkTest.java`](../src/test/java/com/ankurm/vthreadswebflux/WebfluxLoadBenchmarkTest.java).
Transcripts: [`docs/output/01-io-bound-webflux.txt`](output/01-io-bound-webflux.txt), [`docs/output/02-cpu-bound-webflux.txt`](output/02-cpu-bound-webflux.txt), [`docs/output/03-event-loop-starvation.txt`](output/03-event-loop-starvation.txt), [`docs/output/04-stream-backpressure.txt`](output/04-stream-backpressure.txt).
## Why this module exists, and why it's separate from `../virtual-threads-benchmark`
The companion post for this module is the three-way comparison: platform threads vs virtual
threads vs WebFlux. The platform-vs-virtual half of that comparison already has a real,
re-run benchmark in [`../virtual-threads-benchmark`](../virtual-threads-benchmark), built for
a different ankurm.com post and reused here rather than duplicated. This module adds the
missing WebFlux leg, measured with the identical client-side load generator
(`java.net.http.HttpClient` backed by a virtual-thread executor, used only as the *client*) so
all three legs come from the same method on the same 2 vCPU sandbox.
It is its own Maven module rather than a third profile inside `../virtual-threads-benchmark`
because `spring-boot-starter-web` (Tomcat) and `spring-boot-starter-webflux` (Netty) on the
same classpath fight over `WebApplicationType.deduceFromClasspath()` -- exactly the kind of
fragile setup a benchmark should not carry. A module with only `spring-boot-starter-webflux`
needs no such workaround.
## A methodology correction made while building this module: single trials are not reliable here
While measuring the I/O-bound endpoint, identical back-to-back single trials at 600
concurrency swung from 795ms to 1247ms wall time on this shared sandbox -- a spread larger
than the actual gap this benchmark exists to measure. A single HTTP load-test run on a
noisy, shared 2 vCPU box is not precise enough to support a "X% faster" claim between two
models that are actually close; it only looks precise because it produces one number.
The fix applied to the I/O-bound scenario and the event-loop-starvation scenario below:
run the load multiple independent times and report the **median** across trials, not a
single shot. This is a real methodology change, not cosmetic -- it changed which model's
number looked better in earlier drafts of this module before the fix was applied.
**This has one important consequence for how to read the numbers below.** The
platform-thread and virtual-thread numbers reused from
[`../virtual-threads-benchmark`](../virtual-threads-benchmark) are **single trials**, measured
for a different post before this variance was discovered there. The WebFlux numbers in this
module are **medians of 5 (I/O) or 3 (event-loop starvation) trials**. Comparing a
denoised median against a single trial is not perfectly apples-to-apples: a precise
percentage gap between WebFlux and virtual threads should be read as directional, not as a
figure you could reproduce to the point. A gap wide enough to swamp the observed ~50%
single-trial swing -- which is the case for every comparison against platform threads in
this post, and turned out to be the case for WebFlux vs. virtual threads too once
de-noised -- is the part worth trusting.
## I/O-bound: `/io`, `Mono.delay(300ms)`, concurrency 600, median of 5 trials
```
webflux (netty) : total=600 success=600 wall=569ms p50=455ms p99=510ms
```
`Mono.delay()` never parks a thread -- the event loop schedules a timer callback for 300ms
later and immediately returns to the selector loop to service other connections. Against the
single-trial platform/virtual-thread numbers in
[`../virtual-threads-benchmark/docs/output/01-io-bound-benchmark.txt`](../virtual-threads-benchmark/docs/output/01-io-bound-benchmark.txt)
(`platform wall=1325ms p50=733ms p99=1201ms`, `virtual wall=1042ms p50=658ms p99=720ms`),
WebFlux's median-of-5 number is faster on all three metrics: roughly 55-58% faster than
platform threads and 29-45% faster than virtual threads, depending on the metric. Platform
threads being clearly worse is a robust finding -- even the noisiest single WebFlux trial
observed while building this module (1247ms) still beats platform threads' 1325ms. The
WebFlux-vs-virtual-threads gap specifically should be read with the single-trial-vs-median
caveat above in mind, and its exact size moved by roughly 5-10% between the last two
verification runs of this same median-of-5 measurement -- both models avoid Tomcat's
thread-pool queueing entirely and are dramatically faster than platform threads for this
workload, which is the reproducible part of this result.
## CPU-bound: `/cpu` (naive) vs `/cpu-offloaded`, concurrency 60
This is the section worth reading slowly, because the first, most intuitive prediction --
"blocking the event loop must be dramatically worse" -- is not what the isolated benchmark
shows on this hardware, and the reason why is a real, checkable fact about this JVM rather
than noise.
```
LoopResources.DEFAULT_IO_WORKER_COUNT (Netty event-loop threads) = 4
Runtime.availableProcessors() (Schedulers.parallel() thread count) = 2
webflux naive (Mono.fromCallable, no subscribeOn) : total=60 success=60 wall=302ms p50=145ms p99=288ms
webflux offloaded (subscribeOn(Schedulers.parallel())) : total=60 success=60 wall=271ms p50=142ms p99=261ms
```
`/cpu` wraps the identical 20,000x SHA-256 loop from
[`../virtual-threads-benchmark`'s `DemoController`](../virtual-threads-benchmark/src/main/java/com/ankurm/vthreads/DemoController.java)
in `Mono.fromCallable()` with no `subscribeOn(...)` -- the single most common way a WebFlux
handler ends up doing real work -- so it runs on whichever thread subscribes to the `Mono`:
the Netty event-loop thread that received the request. `/cpu-offloaded` moves the identical
work to `Schedulers.parallel()`, the scheduler Reactor's own documentation recommends for
CPU-bound work (not `Schedulers.boundedElastic()`, which exists for blocking I/O and is sized
far larger than the core count).
<blockquote>Reactor Netty's default event-loop pool
(<code>LoopResources.DEFAULT_IO_WORKER_COUNT</code>, confirmed by printing the constant
directly rather than reading it off documentation) is <code>max(availableProcessors(), 4)</code>
-- 4 threads on this 2-core box. <code>Schedulers.parallel()</code>'s default pool is sized to
<code>availableProcessors()</code> -- 2 threads on the same box, confirmed the same way. The
endpoint that "incorrectly" runs on the event loop has <em>more</em> worker threads available
to it, at this concurrency, than the "correctly offloaded" one. That is why offloaded is only
roughly 2-10% faster across these three metrics rather than showing a dramatic gap -- not
because the advice to offload CPU work is wrong.</blockquote>
Offloaded is faster here on every metric, consistent with the advice, but the margin alone
understates why the naive version is still wrong. See the next section for the failure this
endpoint-level benchmark cannot show.
## What the isolated CPU benchmark hides: event-loop starvation of *other* traffic
```
/io alone (baseline, no concurrent CPU load) : total=20 success=20 wall=347ms p50=332ms p99=346ms
/io while 150x /cpu (naive) run concurrently : total=20 success=20 wall=995ms p50=742ms p99=950ms
/io while 150x /cpu-offloaded run concurrently : total=20 success=20 wall=731ms p50=693ms p99=697ms
```
(median of 3 trials each; see the methodology correction above for why)
This is the real cost of the naive endpoint, and benchmarking `/cpu` by itself cannot show it:
`/io` and `/cpu` share the same small event-loop pool. Firing concurrent `/cpu` requests and,
while they are still in flight, firing 20 unrelated `/io` requests at the same server
measures what happens to traffic that has nothing to do with the CPU-bound endpoint.
**Getting a reproducible number here took two fixes, not one.** This test's first cut used
only 8 concurrent `/cpu` requests and produced an inconsistent, sometimes-inverted result
across repeated runs: 8 requests drain through 4 event-loop threads in two short rounds,
finishing well before the `/io` measurement window was over. Raising the load to 60
concurrent requests (matching the CPU-bound benchmark's own concurrency) fixed the
under-loading problem but still left a margin thin enough to flake once during verification
(372ms vs 374ms p99 -- a real tie, not a real result). The final fix was **150 concurrent
`/cpu` requests, median of 3 trials** -- both changes, not a threshold that happened to pass
once.
**A second, more interesting finding came out of raising the load this high**: at 150
concurrent CPU-bound requests -- far beyond the 2 physical cores available -- *even the
offloaded case* now degrades `/io` noticeably (baseline p99 346ms vs. offloaded-load p99
697ms, roughly 2x). Offloading moves the CPU work off the event loop's 4 threads onto
`Schedulers.parallel()`'s 2 threads, which stops it from directly starving `/io`'s access to
the event loop -- but it does not stop it from saturating the 2 physical cores those event-loop
threads still need CPU time on. Naive is still clearly worse than offloaded (roughly 7-36%
worse across wall/p50/p99, varying by metric and by trial), which is the mechanism this test
is built to demonstrate; it is just not a free pass to "offloaded means no impact at all" once
concurrent CPU load exceeds what the box can actually run at once. At higher production
concurrency, both the direct mechanism (naive holding event-loop threads) and this indirect
one (any CPU-bound load competing for the same physical cores) matter.
<ul>
<li>If you're auditing your own WebFlux service: grep for <code>Mono.fromCallable</code>,
<code>Flux.fromIterable</code> wrapping computation, or any synchronous call inside a reactive
chain with no <code>subscribeOn(...)</code> after it -- each one runs on the event loop by
default.</li>
<li><code>Schedulers.parallel()</code> for CPU-bound work, <code>Schedulers.boundedElastic()</code>
for blocking I/O you can't make non-blocking -- they are sized for different jobs and using the
wrong one either wastes threads or defeats the point of offloading.</li>
<li>Reactor Netty event-loop sizing: <a href="https://projectreactor.io/docs/netty/release/reference/index.html#_event_loop_workers" rel="nofollow">projectreactor.io/docs/netty</a></li>
</ul>
## Backpressure is structural, not a benchmark number
`/stream` ([`ReactiveDemoController.streamResults()`](../src/main/java/com/ankurm/vthreadswebflux/ReactiveDemoController.java))
and [`StreamBackpressureTest`](../src/test/java/com/ankurm/vthreadswebflux/StreamBackpressureTest.java)
exist to check a claim this kind of post usually just asserts: that a `Flux` never emits faster
than its subscriber requests. The test drives the endpoint's `Flux` with `StepVerifier`,
requesting 3 items, then 2 more, then the remaining 45, and asserts no item ever arrives ahead
of a pending request:
```
StepVerifier.create(controller.streamResults(), 3)
.expectNext("event-0", "event-1", "event-2")
.expectNoEvent(Duration.ofMillis(80)) -- no 4th item arrives without a request
.thenRequest(2).expectNext("event-3", "event-4")
.thenRequest(45).expectNextCount(45)
.expectComplete()
```
(`docs/output/04-stream-backpressure.txt`) The endpoint is built on `Flux.range()` rather than
the more "realistic-looking" `Flux.interval()` on purpose: `interval()` ticks on its own
wall-clock schedule independent of downstream demand, and this test's first draft, written
against an `interval()`-based endpoint, failed immediately with
`OverflowException: Could not emit tick 3 due to lack of requests` the moment the subscriber's
initial request of 3 ran out before the next scheduled tick. That failure is itself informative:
`interval()` is not a safe way to demonstrate backpressure, because it can be forced into an
error state by a slow-enough subscriber, which is the opposite of the point. `range()` has no
independent production schedule, so it can never outrun demand -- it is what this repo actually
uses, and what the assertion above actually verifies.
Back to [README](../README.md).
@@ -0,0 +1,12 @@
I/O-bound endpoint (/io, Mono.delay(300ms)), concurrency=600, median of 5 trials, 2 vCPU sandbox, Spring Boot 4.1.1 / JDK 25.0.4.1
==================================================================================================================================
webflux (netty) : total=600 success=600 wall=569ms p50=455ms p99=510ms
Mono.delay() never parks a thread -- the event loop schedules a timer callback and
goes back to the selector loop immediately. A single trial at this concurrency swung
50%+ between back-to-back runs on this shared sandbox -- bigger than the gap being
measured -- so the number above is the median of 5 independent trials, not
one shot. Even so it lands in the same range as virtual threads' /io result in
../virtual-threads-benchmark/docs/output/01-io-bound-benchmark.txt (a single trial):
treat any single-run percentage gap between WebFlux and virtual threads here as
noise-level, not a reliable ranking.
@@ -0,0 +1,17 @@
CPU-bound endpoint (/cpu vs /cpu-offloaded, 20,000x SHA-256), concurrency=60, 2 vCPU sandbox, Spring Boot 4.1.1 / JDK 25.0.4.1
==============================================================================================================================
LoopResources.DEFAULT_IO_WORKER_COUNT (Netty event-loop threads) = 4
Runtime.availableProcessors() (Schedulers.parallel() thread count) = 2
webflux naive (Mono.fromCallable, no subscribeOn) : total=60 success=60 wall=302ms p50=145ms p99=288ms
webflux offloaded (subscribeOn(Schedulers.parallel())) : total=60 success=60 wall=271ms p50=142ms p99=261ms
Counter-intuitive result, and worth stating honestly rather than forcing the expected
story: on THIS box, the two wall times are close, because Reactor Netty's default event-
loop pool (DEFAULT_IO_WORKER_COUNT = max(availableProcessors(), 4) = 4 here) is actually
LARGER than Schedulers.parallel()'s pool (sized to availableProcessors() = 2). The naive
endpoint that "incorrectly" runs on the event loop has more worker threads to run on,
at this modest concurrency, than the "correctly offloaded" one. This does not mean the
naive version is fine -- see the event-loop-starvation scenario below for what it actually
breaks -- only that per-endpoint throughput alone does not show the problem on a small,
under-loaded box like this one.
@@ -0,0 +1,19 @@
/io latency while 150 concurrent CPU-bound requests run, median of 3 trials, 2 vCPU sandbox, Spring Boot 4.1.1 / JDK 25.0.4.1
=============================================================================================================================
/io alone (baseline, no concurrent CPU load) : total=20 success=20 wall=347ms p50=332ms p99=346ms
/io while 150x /cpu (naive) run concurrently : total=20 success=20 wall=995ms p50=742ms p99=950ms
/io while 150x /cpu-offloaded run concurrently : total=20 success=20 wall=731ms p50=693ms p99=697ms
This is the real cost of the naive endpoint, and it does not show up by benchmarking
/cpu in isolation: /io shares the same small event-loop pool with /cpu. Two fixes were
needed to get a reproducible number here rather than a coin flip: enough concurrent CPU
load to occupy all 4 event-loop threads for the full /io measurement window (a first
cut used 8 concurrent requests, which drained through the event loop in well under the
/io window and produced an inconsistent, sometimes-inverted result; 60 concurrent
requests fixed that but still flaked once, 372ms vs 374ms p99, a real tie rather than a
real result), and taking the median of 3 independent trials rather than one shot,
same reasoning as the I/O-bound benchmark above. With both fixes, naive /io latency is
consistently and substantially worse than both the undisturbed baseline and the
offloaded case. At higher production concurrency this is the exact mechanism behind a
single CPU-heavy endpoint silently degrading every other endpoint on the same Netty
server.
@@ -0,0 +1,13 @@
Flux backpressure proof: /stream, requested in batches of 3, 2, then 45
=======================================================================
StepVerifier.create(controller.streamResults(), 3)
.expectNext("event-0", "event-1", "event-2")
.expectNoEvent(Duration.ofMillis(80)) -- no 4th item arrives without a request
.thenRequest(2).expectNext("event-3", "event-4")
.thenRequest(45).expectNextCount(45)
.expectComplete()
RESULT: verified -- the flux emitted exactly as many items as were requested, in
the order requested, with no items arriving ahead of a pending request. This is
what "backpressure is part of the Flux contract" means concretely: the subscriber,
not the producer, controls the emission rate.