1
0
Files
vector-api-jep-537-demo/README.md
asmhatre 780848e102 Complete runnable Vector API (JEP 537) demo
SIMD kernels (square-and-add, masked filter, brightness clamp, FP reduction ordering), an exploratory timing harness, JMH benchmarks with Maven and Maven-free builds, run scripts, README and the captured results with their exact environment.
2026-07-26 11:50:15 +00:00

4.9 KiB

Java Vector API (JEP 537) — runnable demo

Companion code for Java Vector API (JEP 537): SIMD, Auto-Vectorization, Masks, and Real Benchmarks on ankurm.com. Every code listing and every number in that article comes from this repository.

What this is about

Your CPU has vector registers — 256 bits wide on AVX2, 512 on AVX-512, 128 on Arm NEON — that hold several numbers side by side, in lanes. A single SIMD instruction applies one operation to every lane at once. A plain Java for loop touches one element per iteration, so most of that hardware sits idle.

Java gives you two ways to use it:

  • Auto-vectorization (free, implicit). HotSpot's C2 compiler has a SuperWord pass that fuses consecutive scalar iterations into vector ones. When it fires you get SIMD for nothing. It fires only when C2 can prove the transform is safe and profitable — no data-dependent branches, no aliasing it cannot rule out, regular access patterns — and when it declines, it does so silently.
  • The Vector API (explicit). jdk.incubator.vector lets you state the vector intent yourself: load a lane-width slice, do lane-wise arithmetic, store it back, and use masks instead of branches. It does not promise a fixed instruction sequence on every CPU, but it gives far more predictable access to SIMD than hoping C2 infers it.

The API is still incubating — JEP 537 is its twelfth incubation, targeted at JDK 27 — because it is waiting on Project Valhalla's value classes. The shape of the API is stable; the package name will change when it is promoted.

This repo demonstrates both paths, and — more usefully — shows how to tell which one you are on, by running the same kernels with -XX:-UseSuperWord.

Layout

timings/    exploratory timing experiments (plain javac, no dependencies)
  src/main/java/com/ankurm/vectorapi/
    Main.java            runs everything, prints the environment first
    SquareAdd.java       c[i] = -(a[i]^2 + b[i]^2) — the loop C2 vectorizes for free
    MaskedFilter.java    sum elements above a threshold — masks vs a branch
    Brightness.java      scale + bias + clamp — max/min vs two branches
    ReductionOrder.java  why vector and scalar sums differ in the last bits
    Bench.java           the (deliberately simple) timing harness
    Env.java, Data.java, Sink.java
  run.sh

jmh/        the same kernels under JMH — the numbers worth defending
  src/main/java/com/ankurm/vectorapi/
    VectorBench.java     square-and-add, incl. a cautionary checksum variant
    MaskedBench.java     masked filter-and-sum
    BrightnessBench.java brightness + clamp
  pom.xml
  run.sh                 Maven build + run
  run-without-maven.sh   same thing with curl + javac, if you have no Maven

RESULTS.md  everything the article quotes, with the exact machine it came from

Requirements

JDK 21 or newer (anything with jdk.incubator.vector). Maven only for jmh/run.sh. The incubator module must be named explicitly at compile time and run time:

javac --add-modules jdk.incubator.vector ...
java  --add-modules jdk.incubator.vector ...

Running it

# 1. exploratory timings, C2 auto-vectorization on (the default)
./timings/run.sh

# 2. the same experiments with SuperWord disabled — the interesting comparison
./timings/run.sh -XX:-UseSuperWord

# 3. JMH, all benchmarks
./jmh/run.sh

# ...or one class, with a short config
./jmh/run.sh MaskedBench -wi 3 -i 5 -w 1 -r 1 -f 1

timings/run.sh prints the JDK build, OS, CPU count and the preferred vector species before it measures anything, so a pasted result always carries the environment it came from.

What to expect

On the AVX2 machine described in RESULTS.md:

Kernel Scalar Vector API Ratio
square-and-add (JMH) 838.6 ± 106.7 ns 814.5 ± 10.8 ns indistinguishable — C2 already vectorized it
masked filter-and-sum (JMH) 5189.9 ± 250.6 ns 1036.4 ± 27.7 ns 5.0x
brightness + clamp (JMH) 6510.1 ± 302.6 ns 1133.1 ± 17.8 ns 5.7x

The pattern is the point: where C2 auto-vectorizes, explicit code buys nothing; where a data-dependent branch stops it, masks are worth roughly 5x. Confirm which case you are in with -XX:-UseSuperWord before rewriting anything — if the scalar loop does not get slower with that flag, C2 was not vectorizing it in the first place.

These ratios are specific to this JDK, CPU, data size and threshold distribution. Run them on your own hardware before quoting them.

Further reading

License

MIT — see LICENSE.