AtomicLong/LongAdder/VarHandle counters benchmarked against the synchronized and ReentrantLock baselines from the locks module, a deterministic ABA race against a hand-rolled Treiber stack plus the AtomicStampedReference fix, and a VarHandle access-modes demo (plain/opaque/acquire-release/volatile). Co-Authored-By: Claude Sonnet 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01FhzLY5p6okFva3qsnsRyvM
atomics
Companion code for the ankurm.com post "Java Atomics and VarHandle: CAS, LongAdder, and When
Atomics Beat Locks." Third module in java-core-examples, the Java-core / concurrency series.
Versions this was built and tested against
| Component | Version | Notes |
|---|---|---|
| JDK | 25.0.4.1+1 (Temurin, LTS) | |
| JMH | 1.37 | |
| JUnit Jupiter | 5.11.0 | Correctness tests only. |
| Maven | 3.9.11 | |
| Hardware | 2 vCPU x86-64 VM | Same sandbox as jmm and locks - see the honesty note below. |
Quickstart
export JAVA_HOME=/path/to/jdk-25
mvn package
java -jar target/benchmarks.jar IncrementBenchmark -t 8
java -cp target/classes com.ankurm.atomics.AbaProblemDemo
java -cp target/classes com.ankurm.atomics.VarHandleAccessModesDemo
scripts/run-all.sh regenerates every file in output/. scripts/run.sh <BenchmarkClass> [args]
runs one benchmark ad hoc.
What's in here
| File | What it shows |
|---|---|
src/main/java/.../Counter.java |
The interface all five counters implement. |
.../SynchronizedCounter.java, .../ReentrantLockCounter.java |
The two lock-based baselines from the locks module, reused here for a direct comparison. |
.../AtomicLongCounter.java |
AtomicLong.incrementAndGet() - one contended CAS location. |
.../LongAdderCounter.java |
LongAdder - writes spread across striped cells, sum() on read. |
.../VarHandleCounter.java |
The same CAS loop AtomicLong does internally, written by hand against a plain field via VarHandle. |
.../TreiberStack.java |
A textbook lock-free stack - and where the ABA problem actually lives. |
.../AbaProblemDemo.java |
A real, two-thread, deterministically-forced ABA race against TreiberStack, plus the AtomicStampedReference fix. |
.../VarHandleAccessModesDemo.java |
VarHandle's four access-mode families (plain, opaque, acquire/release, volatile) plus compareAndSet. |
src/test/java/.../CounterCorrectnessTest.java |
Lost-update sanity checks - does NOT prove throughput, ABA, or ordering claims, see its Javadoc. |
output/01 |
IncrementBenchmark swept across 1, 2, 4, 8, 16, 32, 64 threads. |
output/02 |
The ABA demo's actual output - a real corruption, then the stamped fix. |
output/03 |
The VarHandle access-modes demo's output. |
output/04 |
JUnit correctness run. |
Reading the numbers honestly (2-vCPU sandbox)
LongAdder wins under any real contention (output/01): from 2 threads onward it holds
~160,000-177,000 ops/ms, 3-4x every other counter, exactly matching its design - writes land on
one of several striped cells instead of fighting over one location, so contention on the counter
itself mostly disappears. At 1 thread it's actually the slowest of the five (no contention to
amortize the cell-array bookkeeping against), which is the one case its own Javadoc explicitly
says to expect.
AtomicLong is fastest at 1 thread, then drops hard and plateaus (output/01): ~152,000
ops/ms uncontended, down to ~43,000-58,000 from 2 threads on - every thread past the first is
now genuinely fighting over one CAS location, and on this 2-core box that settles into a stable
but much lower plateau rather than degrading further.
VarHandle's hand-written CAS loop is consistently behind AtomicLong's under contention -
roughly 30,000-32,000 ops/ms against AtomicLong's 43,000-58,000 from 2 threads onward, despite
both doing the same fundamental operation (read, compute, compareAndSet, retry on failure).
This repository does not have a definitive answer for the exact gap; the likely explanation is
that AtomicLong.incrementAndGet() is a JIT/JVM intrinsic HotSpot recognizes and compiles
specially, while the hand-written getVolatile + compareAndSet loop in VarHandleCounter,
while using the same underlying CAS instruction, doesn't get quite the same treatment. Take this
as "hand-rolling the loop yourself costs something measurable here," not as a precise multiplier
to expect elsewhere.
synchronized again lags well behind at 2+ threads - the same pattern documented in the
locks module's README, reproduced here with a completely different set of counters, which is
some evidence it's a property of this sandbox's monitor-contention path rather than a fluke of
one benchmark file.
License
MIT - see the repo-wide LICENSE.