Files
asmhatreandClaude Sonnet 5 ff183da163 atomics: Java Atomics and VarHandle companion code
AtomicLong/LongAdder/VarHandle counters benchmarked against the synchronized
and ReentrantLock baselines from the locks module, a deterministic ABA race
against a hand-rolled Treiber stack plus the AtomicStampedReference fix, and
a VarHandle access-modes demo (plain/opaque/acquire-release/volatile).

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FhzLY5p6okFva3qsnsRyvM
2026-09-30 06:23:03 +00:00
..

atomics

Companion code for the ankurm.com post "Java Atomics and VarHandle: CAS, LongAdder, and When Atomics Beat Locks." Third module in java-core-examples, the Java-core / concurrency series.

Versions this was built and tested against

Component Version Notes
JDK 25.0.4.1+1 (Temurin, LTS)
JMH 1.37
JUnit Jupiter 5.11.0 Correctness tests only.
Maven 3.9.11
Hardware 2 vCPU x86-64 VM Same sandbox as jmm and locks - see the honesty note below.

Quickstart

export JAVA_HOME=/path/to/jdk-25
mvn package
java -jar target/benchmarks.jar IncrementBenchmark -t 8
java -cp target/classes com.ankurm.atomics.AbaProblemDemo
java -cp target/classes com.ankurm.atomics.VarHandleAccessModesDemo

scripts/run-all.sh regenerates every file in output/. scripts/run.sh <BenchmarkClass> [args] runs one benchmark ad hoc.

What's in here

File What it shows
src/main/java/.../Counter.java The interface all five counters implement.
.../SynchronizedCounter.java, .../ReentrantLockCounter.java The two lock-based baselines from the locks module, reused here for a direct comparison.
.../AtomicLongCounter.java AtomicLong.incrementAndGet() - one contended CAS location.
.../LongAdderCounter.java LongAdder - writes spread across striped cells, sum() on read.
.../VarHandleCounter.java The same CAS loop AtomicLong does internally, written by hand against a plain field via VarHandle.
.../TreiberStack.java A textbook lock-free stack - and where the ABA problem actually lives.
.../AbaProblemDemo.java A real, two-thread, deterministically-forced ABA race against TreiberStack, plus the AtomicStampedReference fix.
.../VarHandleAccessModesDemo.java VarHandle's four access-mode families (plain, opaque, acquire/release, volatile) plus compareAndSet.
src/test/java/.../CounterCorrectnessTest.java Lost-update sanity checks - does NOT prove throughput, ABA, or ordering claims, see its Javadoc.
output/01 IncrementBenchmark swept across 1, 2, 4, 8, 16, 32, 64 threads.
output/02 The ABA demo's actual output - a real corruption, then the stamped fix.
output/03 The VarHandle access-modes demo's output.
output/04 JUnit correctness run.

Reading the numbers honestly (2-vCPU sandbox)

LongAdder wins under any real contention (output/01): from 2 threads onward it holds ~160,000-177,000 ops/ms, 3-4x every other counter, exactly matching its design - writes land on one of several striped cells instead of fighting over one location, so contention on the counter itself mostly disappears. At 1 thread it's actually the slowest of the five (no contention to amortize the cell-array bookkeeping against), which is the one case its own Javadoc explicitly says to expect.

AtomicLong is fastest at 1 thread, then drops hard and plateaus (output/01): ~152,000 ops/ms uncontended, down to ~43,000-58,000 from 2 threads on - every thread past the first is now genuinely fighting over one CAS location, and on this 2-core box that settles into a stable but much lower plateau rather than degrading further.

VarHandle's hand-written CAS loop is consistently behind AtomicLong's under contention - roughly 30,000-32,000 ops/ms against AtomicLong's 43,000-58,000 from 2 threads onward, despite both doing the same fundamental operation (read, compute, compareAndSet, retry on failure). This repository does not have a definitive answer for the exact gap; the likely explanation is that AtomicLong.incrementAndGet() is a JIT/JVM intrinsic HotSpot recognizes and compiles specially, while the hand-written getVolatile + compareAndSet loop in VarHandleCounter, while using the same underlying CAS instruction, doesn't get quite the same treatment. Take this as "hand-rolling the loop yourself costs something measurable here," not as a precise multiplier to expect elsewhere.

synchronized again lags well behind at 2+ threads - the same pattern documented in the locks module's README, reproduced here with a completely different set of counters, which is some evidence it's a property of this sandbox's monitor-contention path rather than a fluke of one benchmark file.

License

MIT - see the repo-wide LICENSE.