# Captured results Everything below was produced by the code in this repository. Your numbers will differ — that is the point of shipping the code rather than only the table. ## Environment ``` JDK : OpenJDK 21.0.2+13-58 (HotSpot 64-Bit Server VM, mixed mode) OS : Ubuntu 22.04.5 LTS, Linux 6.8.0-124-generic, x86_64 CPU : AMD Ryzen 5 5600U (Zen 3), 2 vCPUs (1 core / 2 threads) exposed Hypervisor : Microsoft Hyper-V, full virtualization ISA : AVX2 + FMA; no AVX-512 -> FloatVector.SPECIES_PREFERRED = 8 lanes, 256-bit Caches : L1d 32 KiB, L2 512 KiB, L3 16 MiB (as reported to the guest) JVM flags : --add-modules jdk.incubator.vector (plus -XX:-UseSuperWord where noted) Threads : single-threaded throughout Data : 8192 floats = 32 KiB per array, cache-resident Frequency : not controllable inside the guest (no cpufreq governor exposed) ``` The last line matters: on a virtualized, frequency-scaled machine, absolute nanoseconds are soft. Ratios measured back-to-back in the same JVM are far more stable than the absolute numbers, and the JMH table is more trustworthy than the exploratory one. ## 1. Exploratory timings — `timings/run.sh` Harness: 50,000 warmup calls, then 15 rounds of 5,000 calls; median reported. These are teaching experiments, not statistically rigorous benchmarks. ``` --- kernel 1: c[i] = -(a[i]^2 + b[i]^2), 8192 floats --- scalar for-loop median 832.0 ns/call best 819.6 ns/call Vector API median 1279.9 ns/call best 1207.8 ns/call vector advantage 0.65x --- kernel 2: masked filter-and-sum, 8192 floats --- scalar sum = 3084.1404 vector sum = 3084.1448 (same value, different rounding) scalar (branch) median 5021.9 ns/call best 4926.9 ns/call Vector API (masked) median 1029.9 ns/call best 1021.4 ns/call vector advantage 4.88x --- kernel 3: image brightness + clamp, 8192 pixels --- scalar (branchy clamp) median 6881.6 ns/call best 6839.2 ns/call Vector API (max/min) median 1243.8 ns/call best 1222.6 ns/call vector advantage 5.53x --- floating-point reduction ordering --- scalar (left to right) : 4103.688965 vector (lanes, then reduce): 4103.681641 ``` ## 2. Same experiments with auto-vectorization off — `timings/run.sh -XX:-UseSuperWord` ``` --- kernel 1 --- scalar 3427.0 ns/call vector 1230.1 ns/call -> 2.79x --- kernel 2 --- scalar 5040.4 ns/call vector 1031.3 ns/call -> 4.89x --- kernel 3 --- scalar 6883.9 ns/call vector 1364.1 ns/call -> 5.05x ``` Read the two runs together — that comparison is the whole experiment: | Scalar loop | SuperWord ON | SuperWord OFF | Conclusion | |---|---|---|---| | kernel 1 (straight-line math) | 832 ns | 3427 ns | C2 **was** vectorizing it; disabling SuperWord costs 4.1x | | kernel 2 (data-dependent branch) | 5022 ns | 5040 ns | unchanged, so C2 was **not** vectorizing it | | kernel 3 (branchy clamp) | 6882 ns | 6884 ns | unchanged, so C2 was **not** vectorizing it | The Vector API versions are essentially unmoved by the flag in all three cases, which is the property the API is actually selling. Run-to-run spread on this machine was a few percent for kernels 1 and 2, and up to ~13% for kernel 3 (5.05x–6.28x observed across runs) — another reminder to quote JMH rather than a stopwatch. ## 3. JMH — `jmh/run.sh` JMH 1.37, 1 fork, 3 warmup iterations x 1 s, 5 measurement iterations x 1 s, single thread, average time per operation. ``` Benchmark Mode Cnt Score Error Units VectorBench.scalar avgt 5 838.631 ± 106.657 ns/op VectorBench.vector avgt 5 814.531 ± 10.766 ns/op VectorBench.vectorChecksum avgt 5 7409.993 ± 224.753 ns/op MaskedBench.scalarBranch avgt 5 5189.874 ± 250.597 ns/op MaskedBench.vectorMasked avgt 5 1036.419 ± 27.726 ns/op BrightnessBench.scalarClamp avgt 5 6510.091 ± 302.638 ns/op BrightnessBench.vectorClamp avgt 5 1133.073 ± 17.763 ns/op ``` Three readings: 1. **Square-and-add: no measurable difference.** 838.6 ± 106.7 vs 814.5 ± 10.8 — the error bars overlap. On a loop C2 already vectorizes, hand-writing the vector loop bought nothing measurable here. (The exploratory harness made the vector version look ~35% slower; that gap is harness overhead, not the kernel. This is exactly why the JMH number is the one to quote.) 2. **Masked filter: 5.0x** (5189.9 / 1036.4), with tight error bars on the vector side. 3. **Brightness clamp: 5.7x** (6510.1 / 1133.1). `vectorChecksum` is a deliberate cautionary example: returning a checksum does prevent dead-code elimination, but the checksum loop is then *inside* the measurement, and the benchmark reports ~9x the time of the kernel it was meant to measure. Consuming the output array through a `Blackhole` keeps the measurement honest without adding work. ## What is not reproduced here The article shows an excerpt of the AVX2 machine code C2 emits for the square-and-add loop. Reproducing that needs an `hsdis` disassembler plugin installed next to your JDK: ``` java --add-modules jdk.incubator.vector \ -XX:+UnlockDiagnosticVMOptions -XX:+PrintAssembly \ -XX:CompileCommand=print,com.ankurm.vectorapi.SquareAdd::vectorComputation \ -cp out com.ankurm.vectorapi.Main ``` Without `hsdis` the JVM prints `Loading hsdis library failed` and falls back to a hex dump.