Files
spring-async-demo/virtual-threads-benchmark-webflux/docs/output/03-event-loop-starvation.txt
T
Claude 09631dcaab Add virtual-threads-benchmark-webflux: the WebFlux leg of the three-way benchmark
Fixes found during self-correction before publishing:
- /stream used Flux.interval(), which ticks on its own wall-clock schedule
  independent of downstream demand and threw OverflowException under a slow
  subscriber; switched to Flux.range(), which has no independent production
  schedule and can never outrun demand.
- Single-trial HTTP load tests on this shared sandbox swung by more than 50%
  run to run (795ms-1247ms observed on the identical /io endpoint back to
  back) -- large enough to flip which threading model looked faster. Fixed
  by taking the median of 5 independent trials for the I/O-bound benchmark
  and the median of 3 for the event-loop-starvation benchmark, rather than
  reporting a single noisy run as if it were precise.
- The event-loop-starvation test's first cut used only 8 concurrent /cpu
  requests as background load, which drained through the 4 event-loop
  threads well inside the /io measurement window and produced an
  inconsistent, sometimes-inverted result across runs; raising to 60 fixed
  the under-loading problem but still flaked once during verification
  (372ms vs 374ms p99, a real tie). Final fix: 150 concurrent requests plus
  the median-of-3 trials above.

Also adds StreamBackpressureTest, a StepVerifier proof that the /stream
endpoint never emits ahead of its subscriber's outstanding requests, and
updates the module's docs to report the de-noised numbers with an explicit
methodology note on how they compare to the single-trial platform/virtual-
thread numbers reused from a different post.
2026-09-19 09:34:18 +00:00

20 lines
1.6 KiB
Plaintext

/io latency while 150 concurrent CPU-bound requests run, median of 3 trials, 2 vCPU sandbox, Spring Boot 4.1.1 / JDK 25.0.4.1
=============================================================================================================================
/io alone (baseline, no concurrent CPU load) : total=20 success=20 wall=347ms p50=332ms p99=346ms
/io while 150x /cpu (naive) run concurrently : total=20 success=20 wall=995ms p50=742ms p99=950ms
/io while 150x /cpu-offloaded run concurrently : total=20 success=20 wall=731ms p50=693ms p99=697ms
This is the real cost of the naive endpoint, and it does not show up by benchmarking
/cpu in isolation: /io shares the same small event-loop pool with /cpu. Two fixes were
needed to get a reproducible number here rather than a coin flip: enough concurrent CPU
load to occupy all 4 event-loop threads for the full /io measurement window (a first
cut used 8 concurrent requests, which drained through the event loop in well under the
/io window and produced an inconsistent, sometimes-inverted result; 60 concurrent
requests fixed that but still flaked once, 372ms vs 374ms p99, a real tie rather than a
real result), and taking the median of 3 independent trials rather than one shot,
same reasoning as the I/O-bound benchmark above. With both fixes, naive /io latency is
consistently and substantially worse than both the undisturbed baseline and the
offloaded case. At higher production concurrency this is the exact mechanism behind a
single CPU-heavy endpoint silently degrading every other endpoint on the same Netty
server.