Add virtual-threads-benchmark-webflux: the WebFlux leg of the three-way benchmark
Fixes found during self-correction before publishing: - /stream used Flux.interval(), which ticks on its own wall-clock schedule independent of downstream demand and threw OverflowException under a slow subscriber; switched to Flux.range(), which has no independent production schedule and can never outrun demand. - Single-trial HTTP load tests on this shared sandbox swung by more than 50% run to run (795ms-1247ms observed on the identical /io endpoint back to back) -- large enough to flip which threading model looked faster. Fixed by taking the median of 5 independent trials for the I/O-bound benchmark and the median of 3 for the event-loop-starvation benchmark, rather than reporting a single noisy run as if it were precise. - The event-loop-starvation test's first cut used only 8 concurrent /cpu requests as background load, which drained through the 4 event-loop threads well inside the /io measurement window and produced an inconsistent, sometimes-inverted result across runs; raising to 60 fixed the under-loading problem but still flaked once during verification (372ms vs 374ms p99, a real tie). Final fix: 150 concurrent requests plus the median-of-3 trials above. Also adds StreamBackpressureTest, a StepVerifier proof that the /stream endpoint never emits ahead of its subscriber's outstanding requests, and updates the module's docs to report the de-noised numbers with an explicit methodology note on how they compare to the single-trial platform/virtual- thread numbers reused from a different post.
This commit is contained in:
@@ -0,0 +1,19 @@
|
||||
/io latency while 150 concurrent CPU-bound requests run, median of 3 trials, 2 vCPU sandbox, Spring Boot 4.1.1 / JDK 25.0.4.1
|
||||
=============================================================================================================================
|
||||
/io alone (baseline, no concurrent CPU load) : total=20 success=20 wall=347ms p50=332ms p99=346ms
|
||||
/io while 150x /cpu (naive) run concurrently : total=20 success=20 wall=995ms p50=742ms p99=950ms
|
||||
/io while 150x /cpu-offloaded run concurrently : total=20 success=20 wall=731ms p50=693ms p99=697ms
|
||||
|
||||
This is the real cost of the naive endpoint, and it does not show up by benchmarking
|
||||
/cpu in isolation: /io shares the same small event-loop pool with /cpu. Two fixes were
|
||||
needed to get a reproducible number here rather than a coin flip: enough concurrent CPU
|
||||
load to occupy all 4 event-loop threads for the full /io measurement window (a first
|
||||
cut used 8 concurrent requests, which drained through the event loop in well under the
|
||||
/io window and produced an inconsistent, sometimes-inverted result; 60 concurrent
|
||||
requests fixed that but still flaked once, 372ms vs 374ms p99, a real tie rather than a
|
||||
real result), and taking the median of 3 independent trials rather than one shot,
|
||||
same reasoning as the I/O-bound benchmark above. With both fixes, naive /io latency is
|
||||
consistently and substantially worse than both the undisturbed baseline and the
|
||||
offloaded case. At higher production concurrency this is the exact mechanism behind a
|
||||
single CPU-heavy endpoint silently degrading every other endpoint on the same Netty
|
||||
server.
|
||||
Reference in New Issue
Block a user