diff --git a/README.md b/README.md index cf0845b..f9dcf5c 100644 --- a/README.md +++ b/README.md @@ -1,3 +1,117 @@ -# vector-api-jep-537-demo +# Java Vector API (JEP 537) — runnable demo -Runnable companion code for the Java Vector API (JEP 537) guide on ankurm.com: SIMD kernels, masks, the loopBound+tail pattern, and JMH benchmarks on JDK 21. \ No newline at end of file +Companion code for **[Java Vector API (JEP 537): SIMD, Auto-Vectorization, Masks, and Real Benchmarks](https://ankurm.com/java-vector-api-jep-537-simd-guide/)** on ankurm.com. +Every code listing and every number in that article comes from this repository. + +## What this is about + +Your CPU has vector registers — 256 bits wide on AVX2, 512 on AVX-512, 128 on +Arm NEON — that hold several numbers side by side, in *lanes*. A single SIMD +instruction applies one operation to every lane at once. A plain Java `for` loop +touches one element per iteration, so most of that hardware sits idle. + +Java gives you two ways to use it: + +* **Auto-vectorization (free, implicit).** HotSpot's C2 compiler has a SuperWord + pass that fuses consecutive scalar iterations into vector ones. When it fires + you get SIMD for nothing. It fires only when C2 can prove the transform is safe + and profitable — no data-dependent branches, no aliasing it cannot rule out, + regular access patterns — and when it declines, it does so silently. +* **The Vector API (explicit).** `jdk.incubator.vector` lets you state the vector + intent yourself: load a lane-width slice, do lane-wise arithmetic, store it + back, and use *masks* instead of branches. It does not promise a fixed + instruction sequence on every CPU, but it gives far more predictable access to + SIMD than hoping C2 infers it. + +The API is still incubating — JEP 537 is its twelfth incubation, targeted at JDK +27 — because it is waiting on Project Valhalla's value classes. The shape of the +API is stable; the package name will change when it is promoted. + +This repo demonstrates both paths, and — more usefully — shows how to **tell +which one you are on**, by running the same kernels with `-XX:-UseSuperWord`. + +## Layout + +``` +timings/ exploratory timing experiments (plain javac, no dependencies) + src/main/java/com/ankurm/vectorapi/ + Main.java runs everything, prints the environment first + SquareAdd.java c[i] = -(a[i]^2 + b[i]^2) — the loop C2 vectorizes for free + MaskedFilter.java sum elements above a threshold — masks vs a branch + Brightness.java scale + bias + clamp — max/min vs two branches + ReductionOrder.java why vector and scalar sums differ in the last bits + Bench.java the (deliberately simple) timing harness + Env.java, Data.java, Sink.java + run.sh + +jmh/ the same kernels under JMH — the numbers worth defending + src/main/java/com/ankurm/vectorapi/ + VectorBench.java square-and-add, incl. a cautionary checksum variant + MaskedBench.java masked filter-and-sum + BrightnessBench.java brightness + clamp + pom.xml + run.sh Maven build + run + run-without-maven.sh same thing with curl + javac, if you have no Maven + +RESULTS.md everything the article quotes, with the exact machine it came from +``` + +## Requirements + +JDK 21 or newer (anything with `jdk.incubator.vector`). Maven only for `jmh/run.sh`. +The incubator module must be named explicitly at compile time *and* run time: + +```bash +javac --add-modules jdk.incubator.vector ... +java --add-modules jdk.incubator.vector ... +``` + +## Running it + +```bash +# 1. exploratory timings, C2 auto-vectorization on (the default) +./timings/run.sh + +# 2. the same experiments with SuperWord disabled — the interesting comparison +./timings/run.sh -XX:-UseSuperWord + +# 3. JMH, all benchmarks +./jmh/run.sh + +# ...or one class, with a short config +./jmh/run.sh MaskedBench -wi 3 -i 5 -w 1 -r 1 -f 1 +``` + +`timings/run.sh` prints the JDK build, OS, CPU count and the preferred vector +species before it measures anything, so a pasted result always carries the +environment it came from. + +## What to expect + +On the AVX2 machine described in [RESULTS.md](RESULTS.md): + +| Kernel | Scalar | Vector API | Ratio | +|---|---|---|---| +| square-and-add (JMH) | 838.6 ± 106.7 ns | 814.5 ± 10.8 ns | indistinguishable — C2 already vectorized it | +| masked filter-and-sum (JMH) | 5189.9 ± 250.6 ns | 1036.4 ± 27.7 ns | **5.0x** | +| brightness + clamp (JMH) | 6510.1 ± 302.6 ns | 1133.1 ± 17.8 ns | **5.7x** | + +The pattern is the point: where C2 auto-vectorizes, explicit code buys nothing; +where a data-dependent branch stops it, masks are worth roughly 5x. Confirm which +case you are in with `-XX:-UseSuperWord` before rewriting anything — if the scalar +loop does not get slower with that flag, C2 was not vectorizing it in the first +place. + +These ratios are specific to this JDK, CPU, data size and threshold +distribution. Run them on your own hardware before quoting them. + +## Further reading + +* [The article this repo belongs to](https://ankurm.com/java-vector-api-jep-537-simd-guide/) +* [JEP 537: Vector API (Twelfth Incubator)](https://openjdk.org/jeps/537) +* [`jdk.incubator.vector` package docs](https://docs.oracle.com/en/java/javase/21/docs/api/jdk.incubator.vector/jdk/incubator/vector/package-summary.html) +* [JMH](https://github.com/openjdk/jmh) + +## License + +MIT — see [LICENSE](LICENSE).