5.4 KiB
Captured results
Everything below was produced by the code in this repository. Your numbers will differ — that is the point of shipping the code rather than only the table.
Environment
JDK : OpenJDK 21.0.2+13-58 (HotSpot 64-Bit Server VM, mixed mode)
OS : Ubuntu 22.04.5 LTS, Linux 6.8.0-124-generic, x86_64
CPU : AMD Ryzen 5 5600U (Zen 3), 2 vCPUs (1 core / 2 threads) exposed
Hypervisor : Microsoft Hyper-V, full virtualization
ISA : AVX2 + FMA; no AVX-512 -> FloatVector.SPECIES_PREFERRED = 8 lanes, 256-bit
Caches : L1d 32 KiB, L2 512 KiB, L3 16 MiB (as reported to the guest)
JVM flags : --add-modules jdk.incubator.vector (plus -XX:-UseSuperWord where noted)
Threads : single-threaded throughout
Data : 8192 floats = 32 KiB per array, cache-resident
Frequency : not controllable inside the guest (no cpufreq governor exposed)
The last line matters: on a virtualized, frequency-scaled machine, absolute nanoseconds are soft. Ratios measured back-to-back in the same JVM are far more stable than the absolute numbers, and the JMH table is more trustworthy than the exploratory one.
1. Exploratory timings — timings/run.sh
Harness: 50,000 warmup calls, then 15 rounds of 5,000 calls; median reported. These are teaching experiments, not statistically rigorous benchmarks.
--- kernel 1: c[i] = -(a[i]^2 + b[i]^2), 8192 floats ---
scalar for-loop median 832.0 ns/call best 819.6 ns/call
Vector API median 1279.9 ns/call best 1207.8 ns/call
vector advantage 0.65x
--- kernel 2: masked filter-and-sum, 8192 floats ---
scalar sum = 3084.1404
vector sum = 3084.1448 (same value, different rounding)
scalar (branch) median 5021.9 ns/call best 4926.9 ns/call
Vector API (masked) median 1029.9 ns/call best 1021.4 ns/call
vector advantage 4.88x
--- kernel 3: image brightness + clamp, 8192 pixels ---
scalar (branchy clamp) median 6881.6 ns/call best 6839.2 ns/call
Vector API (max/min) median 1243.8 ns/call best 1222.6 ns/call
vector advantage 5.53x
--- floating-point reduction ordering ---
scalar (left to right) : 4103.688965
vector (lanes, then reduce): 4103.681641
2. Same experiments with auto-vectorization off — timings/run.sh -XX:-UseSuperWord
--- kernel 1 --- scalar 3427.0 ns/call vector 1230.1 ns/call -> 2.79x
--- kernel 2 --- scalar 5040.4 ns/call vector 1031.3 ns/call -> 4.89x
--- kernel 3 --- scalar 6883.9 ns/call vector 1364.1 ns/call -> 5.05x
Read the two runs together — that comparison is the whole experiment:
| Scalar loop | SuperWord ON | SuperWord OFF | Conclusion |
|---|---|---|---|
| kernel 1 (straight-line math) | 832 ns | 3427 ns | C2 was vectorizing it; disabling SuperWord costs 4.1x |
| kernel 2 (data-dependent branch) | 5022 ns | 5040 ns | unchanged, so C2 was not vectorizing it |
| kernel 3 (branchy clamp) | 6882 ns | 6884 ns | unchanged, so C2 was not vectorizing it |
The Vector API versions are essentially unmoved by the flag in all three cases, which is the property the API is actually selling.
Run-to-run spread on this machine was a few percent for kernels 1 and 2, and up to ~13% for kernel 3 (5.05x–6.28x observed across runs) — another reminder to quote JMH rather than a stopwatch.
3. JMH — jmh/run.sh
JMH 1.37, 1 fork, 3 warmup iterations x 1 s, 5 measurement iterations x 1 s, single thread, average time per operation.
Benchmark Mode Cnt Score Error Units
VectorBench.scalar avgt 5 838.631 ± 106.657 ns/op
VectorBench.vector avgt 5 814.531 ± 10.766 ns/op
VectorBench.vectorChecksum avgt 5 7409.993 ± 224.753 ns/op
MaskedBench.scalarBranch avgt 5 5189.874 ± 250.597 ns/op
MaskedBench.vectorMasked avgt 5 1036.419 ± 27.726 ns/op
BrightnessBench.scalarClamp avgt 5 6510.091 ± 302.638 ns/op
BrightnessBench.vectorClamp avgt 5 1133.073 ± 17.763 ns/op
Three readings:
- Square-and-add: no measurable difference. 838.6 ± 106.7 vs 814.5 ± 10.8 — the error bars overlap. On a loop C2 already vectorizes, hand-writing the vector loop bought nothing measurable here. (The exploratory harness made the vector version look ~35% slower; that gap is harness overhead, not the kernel. This is exactly why the JMH number is the one to quote.)
- Masked filter: 5.0x (5189.9 / 1036.4), with tight error bars on the vector side.
- Brightness clamp: 5.7x (6510.1 / 1133.1).
vectorChecksum is a deliberate cautionary example: returning a checksum does
prevent dead-code elimination, but the checksum loop is then inside the
measurement, and the benchmark reports ~9x the time of the kernel it was meant to
measure. Consuming the output array through a Blackhole keeps the measurement
honest without adding work.
What is not reproduced here
The article shows an excerpt of the AVX2 machine code C2 emits for the
square-and-add loop. Reproducing that needs an hsdis disassembler plugin
installed next to your JDK:
java --add-modules jdk.incubator.vector \
-XX:+UnlockDiagnosticVMOptions -XX:+PrintAssembly \
-XX:CompileCommand=print,com.ankurm.vectorapi.SquareAdd::vectorComputation \
-cp out com.ankurm.vectorapi.Main
Without hsdis the JVM prints Loading hsdis library failed and falls back to a
hex dump.