Files
Claude 09631dcaab Add virtual-threads-benchmark-webflux: the WebFlux leg of the three-way benchmark
Fixes found during self-correction before publishing:
- /stream used Flux.interval(), which ticks on its own wall-clock schedule
  independent of downstream demand and threw OverflowException under a slow
  subscriber; switched to Flux.range(), which has no independent production
  schedule and can never outrun demand.
- Single-trial HTTP load tests on this shared sandbox swung by more than 50%
  run to run (795ms-1247ms observed on the identical /io endpoint back to
  back) -- large enough to flip which threading model looked faster. Fixed
  by taking the median of 5 independent trials for the I/O-bound benchmark
  and the median of 3 for the event-loop-starvation benchmark, rather than
  reporting a single noisy run as if it were precise.
- The event-loop-starvation test's first cut used only 8 concurrent /cpu
  requests as background load, which drained through the 4 event-loop
  threads well inside the /io measurement window and produced an
  inconsistent, sometimes-inverted result across runs; raising to 60 fixed
  the under-loading problem but still flaked once during verification
  (372ms vs 374ms p99, a real tie). Final fix: 150 concurrent requests plus
  the median-of-3 trials above.

Also adds StreamBackpressureTest, a StepVerifier proof that the /stream
endpoint never emits ahead of its subscriber's outstanding requests, and
updates the module's docs to report the de-noised numbers with an explicit
methodology note on how they compare to the single-trial platform/virtual-
thread numbers reused from a different post.
2026-09-19 09:34:18 +00:00

188 lines
12 KiB
Markdown

# 1. WebFlux benchmark methodology and results
[README](../README.md) | Companion module: [`../virtual-threads-benchmark`](../virtual-threads-benchmark/README.md) (platform threads vs virtual threads)
Source: [`ReactiveDemoController.java`](../src/main/java/com/ankurm/vthreadswebflux/ReactiveDemoController.java), [`CpuWork.java`](../src/main/java/com/ankurm/vthreadswebflux/CpuWork.java).
Test: [`WebfluxLoadBenchmarkTest.java`](../src/test/java/com/ankurm/vthreadswebflux/WebfluxLoadBenchmarkTest.java).
Transcripts: [`docs/output/01-io-bound-webflux.txt`](output/01-io-bound-webflux.txt), [`docs/output/02-cpu-bound-webflux.txt`](output/02-cpu-bound-webflux.txt), [`docs/output/03-event-loop-starvation.txt`](output/03-event-loop-starvation.txt), [`docs/output/04-stream-backpressure.txt`](output/04-stream-backpressure.txt).
## Why this module exists, and why it's separate from `../virtual-threads-benchmark`
The companion post for this module is the three-way comparison: platform threads vs virtual
threads vs WebFlux. The platform-vs-virtual half of that comparison already has a real,
re-run benchmark in [`../virtual-threads-benchmark`](../virtual-threads-benchmark), built for
a different ankurm.com post and reused here rather than duplicated. This module adds the
missing WebFlux leg, measured with the identical client-side load generator
(`java.net.http.HttpClient` backed by a virtual-thread executor, used only as the *client*) so
all three legs come from the same method on the same 2 vCPU sandbox.
It is its own Maven module rather than a third profile inside `../virtual-threads-benchmark`
because `spring-boot-starter-web` (Tomcat) and `spring-boot-starter-webflux` (Netty) on the
same classpath fight over `WebApplicationType.deduceFromClasspath()` -- exactly the kind of
fragile setup a benchmark should not carry. A module with only `spring-boot-starter-webflux`
needs no such workaround.
## A methodology correction made while building this module: single trials are not reliable here
While measuring the I/O-bound endpoint, identical back-to-back single trials at 600
concurrency swung from 795ms to 1247ms wall time on this shared sandbox -- a spread larger
than the actual gap this benchmark exists to measure. A single HTTP load-test run on a
noisy, shared 2 vCPU box is not precise enough to support a "X% faster" claim between two
models that are actually close; it only looks precise because it produces one number.
The fix applied to the I/O-bound scenario and the event-loop-starvation scenario below:
run the load multiple independent times and report the **median** across trials, not a
single shot. This is a real methodology change, not cosmetic -- it changed which model's
number looked better in earlier drafts of this module before the fix was applied.
**This has one important consequence for how to read the numbers below.** The
platform-thread and virtual-thread numbers reused from
[`../virtual-threads-benchmark`](../virtual-threads-benchmark) are **single trials**, measured
for a different post before this variance was discovered there. The WebFlux numbers in this
module are **medians of 5 (I/O) or 3 (event-loop starvation) trials**. Comparing a
denoised median against a single trial is not perfectly apples-to-apples: a precise
percentage gap between WebFlux and virtual threads should be read as directional, not as a
figure you could reproduce to the point. A gap wide enough to swamp the observed ~50%
single-trial swing -- which is the case for every comparison against platform threads in
this post, and turned out to be the case for WebFlux vs. virtual threads too once
de-noised -- is the part worth trusting.
## I/O-bound: `/io`, `Mono.delay(300ms)`, concurrency 600, median of 5 trials
```
webflux (netty) : total=600 success=600 wall=569ms p50=455ms p99=510ms
```
`Mono.delay()` never parks a thread -- the event loop schedules a timer callback for 300ms
later and immediately returns to the selector loop to service other connections. Against the
single-trial platform/virtual-thread numbers in
[`../virtual-threads-benchmark/docs/output/01-io-bound-benchmark.txt`](../virtual-threads-benchmark/docs/output/01-io-bound-benchmark.txt)
(`platform wall=1325ms p50=733ms p99=1201ms`, `virtual wall=1042ms p50=658ms p99=720ms`),
WebFlux's median-of-5 number is faster on all three metrics: roughly 55-58% faster than
platform threads and 29-45% faster than virtual threads, depending on the metric. Platform
threads being clearly worse is a robust finding -- even the noisiest single WebFlux trial
observed while building this module (1247ms) still beats platform threads' 1325ms. The
WebFlux-vs-virtual-threads gap specifically should be read with the single-trial-vs-median
caveat above in mind, and its exact size moved by roughly 5-10% between the last two
verification runs of this same median-of-5 measurement -- both models avoid Tomcat's
thread-pool queueing entirely and are dramatically faster than platform threads for this
workload, which is the reproducible part of this result.
## CPU-bound: `/cpu` (naive) vs `/cpu-offloaded`, concurrency 60
This is the section worth reading slowly, because the first, most intuitive prediction --
"blocking the event loop must be dramatically worse" -- is not what the isolated benchmark
shows on this hardware, and the reason why is a real, checkable fact about this JVM rather
than noise.
```
LoopResources.DEFAULT_IO_WORKER_COUNT (Netty event-loop threads) = 4
Runtime.availableProcessors() (Schedulers.parallel() thread count) = 2
webflux naive (Mono.fromCallable, no subscribeOn) : total=60 success=60 wall=302ms p50=145ms p99=288ms
webflux offloaded (subscribeOn(Schedulers.parallel())) : total=60 success=60 wall=271ms p50=142ms p99=261ms
```
`/cpu` wraps the identical 20,000x SHA-256 loop from
[`../virtual-threads-benchmark`'s `DemoController`](../virtual-threads-benchmark/src/main/java/com/ankurm/vthreads/DemoController.java)
in `Mono.fromCallable()` with no `subscribeOn(...)` -- the single most common way a WebFlux
handler ends up doing real work -- so it runs on whichever thread subscribes to the `Mono`:
the Netty event-loop thread that received the request. `/cpu-offloaded` moves the identical
work to `Schedulers.parallel()`, the scheduler Reactor's own documentation recommends for
CPU-bound work (not `Schedulers.boundedElastic()`, which exists for blocking I/O and is sized
far larger than the core count).
<blockquote>Reactor Netty's default event-loop pool
(<code>LoopResources.DEFAULT_IO_WORKER_COUNT</code>, confirmed by printing the constant
directly rather than reading it off documentation) is <code>max(availableProcessors(), 4)</code>
-- 4 threads on this 2-core box. <code>Schedulers.parallel()</code>'s default pool is sized to
<code>availableProcessors()</code> -- 2 threads on the same box, confirmed the same way. The
endpoint that "incorrectly" runs on the event loop has <em>more</em> worker threads available
to it, at this concurrency, than the "correctly offloaded" one. That is why offloaded is only
roughly 2-10% faster across these three metrics rather than showing a dramatic gap -- not
because the advice to offload CPU work is wrong.</blockquote>
Offloaded is faster here on every metric, consistent with the advice, but the margin alone
understates why the naive version is still wrong. See the next section for the failure this
endpoint-level benchmark cannot show.
## What the isolated CPU benchmark hides: event-loop starvation of *other* traffic
```
/io alone (baseline, no concurrent CPU load) : total=20 success=20 wall=347ms p50=332ms p99=346ms
/io while 150x /cpu (naive) run concurrently : total=20 success=20 wall=995ms p50=742ms p99=950ms
/io while 150x /cpu-offloaded run concurrently : total=20 success=20 wall=731ms p50=693ms p99=697ms
```
(median of 3 trials each; see the methodology correction above for why)
This is the real cost of the naive endpoint, and benchmarking `/cpu` by itself cannot show it:
`/io` and `/cpu` share the same small event-loop pool. Firing concurrent `/cpu` requests and,
while they are still in flight, firing 20 unrelated `/io` requests at the same server
measures what happens to traffic that has nothing to do with the CPU-bound endpoint.
**Getting a reproducible number here took two fixes, not one.** This test's first cut used
only 8 concurrent `/cpu` requests and produced an inconsistent, sometimes-inverted result
across repeated runs: 8 requests drain through 4 event-loop threads in two short rounds,
finishing well before the `/io` measurement window was over. Raising the load to 60
concurrent requests (matching the CPU-bound benchmark's own concurrency) fixed the
under-loading problem but still left a margin thin enough to flake once during verification
(372ms vs 374ms p99 -- a real tie, not a real result). The final fix was **150 concurrent
`/cpu` requests, median of 3 trials** -- both changes, not a threshold that happened to pass
once.
**A second, more interesting finding came out of raising the load this high**: at 150
concurrent CPU-bound requests -- far beyond the 2 physical cores available -- *even the
offloaded case* now degrades `/io` noticeably (baseline p99 346ms vs. offloaded-load p99
697ms, roughly 2x). Offloading moves the CPU work off the event loop's 4 threads onto
`Schedulers.parallel()`'s 2 threads, which stops it from directly starving `/io`'s access to
the event loop -- but it does not stop it from saturating the 2 physical cores those event-loop
threads still need CPU time on. Naive is still clearly worse than offloaded (roughly 7-36%
worse across wall/p50/p99, varying by metric and by trial), which is the mechanism this test
is built to demonstrate; it is just not a free pass to "offloaded means no impact at all" once
concurrent CPU load exceeds what the box can actually run at once. At higher production
concurrency, both the direct mechanism (naive holding event-loop threads) and this indirect
one (any CPU-bound load competing for the same physical cores) matter.
<ul>
<li>If you're auditing your own WebFlux service: grep for <code>Mono.fromCallable</code>,
<code>Flux.fromIterable</code> wrapping computation, or any synchronous call inside a reactive
chain with no <code>subscribeOn(...)</code> after it -- each one runs on the event loop by
default.</li>
<li><code>Schedulers.parallel()</code> for CPU-bound work, <code>Schedulers.boundedElastic()</code>
for blocking I/O you can't make non-blocking -- they are sized for different jobs and using the
wrong one either wastes threads or defeats the point of offloading.</li>
<li>Reactor Netty event-loop sizing: <a href="https://projectreactor.io/docs/netty/release/reference/index.html#_event_loop_workers" rel="nofollow">projectreactor.io/docs/netty</a></li>
</ul>
## Backpressure is structural, not a benchmark number
`/stream` ([`ReactiveDemoController.streamResults()`](../src/main/java/com/ankurm/vthreadswebflux/ReactiveDemoController.java))
and [`StreamBackpressureTest`](../src/test/java/com/ankurm/vthreadswebflux/StreamBackpressureTest.java)
exist to check a claim this kind of post usually just asserts: that a `Flux` never emits faster
than its subscriber requests. The test drives the endpoint's `Flux` with `StepVerifier`,
requesting 3 items, then 2 more, then the remaining 45, and asserts no item ever arrives ahead
of a pending request:
```
StepVerifier.create(controller.streamResults(), 3)
.expectNext("event-0", "event-1", "event-2")
.expectNoEvent(Duration.ofMillis(80)) -- no 4th item arrives without a request
.thenRequest(2).expectNext("event-3", "event-4")
.thenRequest(45).expectNextCount(45)
.expectComplete()
```
(`docs/output/04-stream-backpressure.txt`) The endpoint is built on `Flux.range()` rather than
the more "realistic-looking" `Flux.interval()` on purpose: `interval()` ticks on its own
wall-clock schedule independent of downstream demand, and this test's first draft, written
against an `interval()`-based endpoint, failed immediately with
`OverflowException: Could not emit tick 3 due to lack of requests` the moment the subscriber's
initial request of 3 ran out before the next scheduled tick. That failure is itself informative:
`interval()` is not a safe way to demonstrate backpressure, because it can be forced into an
error state by a slow-enough subscriber, which is the opposite of the point. `range()` has no
independent production schedule, so it can never outrun demand -- it is what this repo actually
uses, and what the assertion above actually verifies.
Back to [README](../README.md).