Skip to main content

Java Atomics and VarHandle: CAS, LongAdder, and When Atomics Beat Locks

AtomicLong, LongAdder, and hand-written VarHandle CAS loops benchmarked against synchronized and ReentrantLock, plus a real, deterministically-forced ABA race against a Treiber stack and the AtomicStampedReference fix.

A lock-free object pool handed the same pooled connection to two different callers at the same instant, and neither caller’s code was wrong. The pool used a classic lock-free stack for its free list – push a connection back when done, pop one when needed, both via compareAndSet, no locks anywhere, exactly the pattern every “atomics are faster than locks” article recommends. The bug took a week to reproduce on purpose, because it depends on a very specific interleaving: a thread has to read the stack’s top, get paused, and only then does the race matter. This article builds that exact bug on purpose – the ABA problem – alongside the more everyday question of when AtomicLong, LongAdder or a hand-written VarHandle loop actually beats a lock at all.
Versions. Tested on JDK 25.0.4.1+1 (Temurin, LTS). JMH 1.37. All benchmarks ran on a 2-vCPU x86-64 virtual machine – the same sandbox as the rest of this series, and it produces one gap in this article’s numbers that this repository can’t fully explain, reported honestly rather than smoothed over.

Five ways to protect one counter

The same Counter interface as the locks module, five implementations: the two lock-based baselines already benchmarked there, plus three that never take a lock at all.
public interface Counter {
    void increment();
    long get();
}
AtomicLong is a hardware compare-and-swap loop, retried until it wins:
public final class AtomicLongCounter implements Counter {
    private final AtomicLong count = new AtomicLong();
    public void increment() { count.incrementAndGet(); }
    public long get() { return count.get(); }
}
Source: AtomicLongCounter.java. LongAdder takes the opposite approach: instead of every thread fighting over one CAS location, writes spread across an internal array of cells that only appears once contention is actually detected, and sum() adds them all up on read:
public final class LongAdderCounter implements Counter {
    private final LongAdder count = new LongAdder();
    public void increment() { count.increment(); }
    public long get() { return count.sum(); }
}
Source: LongAdderCounter.java. And the same CAS loop AtomicLong does internally, written out by hand against a plain field via VarHandle:
public void increment() {
    long current;
    do {
        current = (long) COUNT.getVolatile(this);
    } while (!COUNT.compareAndSet(this, current, current + 1));
}
Source: VarHandleCounter.java. All five pass the same lost-update sanity check – eight threads, fifty thousand increments each, exact final count asserted – in CounterCorrectnessTest.java (output/04).

Going deeper on this section

Where LongAdder wins, and why AtomicLong doesn’t scale forever

AtomicLong and LongAdder solve contention two structurally different ways:
AtomicLong one CAS location thread A thread B thread C every thread retries against the same location LongAdder A -> cell[0] B -> cell[1] C -> cell[2]
$ java -jar target/benchmarks.jar IncrementBenchmark -t 8 -rf text

Benchmark                          Mode  Cnt       Score      Error   Units
IncrementBenchmark.atomicLong     thrpt    5   58036.390 ± 10538.772  ops/ms
IncrementBenchmark.longAdder      thrpt    5  176823.593 ± 19520.176  ops/ms
IncrementBenchmark.reentrantLock  thrpt    5   49528.714 ±  9256.971  ops/ms
IncrementBenchmark.synchronized_  thrpt    5    8933.358 ±  7612.678  ops/ms
IncrementBenchmark.varHandleCas   thrpt    5   31671.036 ±  9740.912  ops/ms
Output: 01-increment-benchmark-sweep.txt, the 8-thread slice of a sweep across 1, 2, 4, 8, 16, 32 and 64 threads. From 2 threads on, LongAdder holds ~160,000-177,000 ops/ms – 3 to 4x every other counter – because writes land on separate cells instead of fighting over one location. AtomicLong tells the opposite story: fastest of all five at 1 thread (~152,000 ops/ms, no contention to pay for), then it drops hard the moment a second thread shows up and settles into a stable ~43,000-58,000 plateau. LongAdder‘s own class documentation predicts exactly this shape: “Under low update contention, the two classes have similar characteristics. But under high contention, expected throughput of this class is significantly higher, at the expense of higher space consumption.”
The trade-off LongAdder’s number hides. sum() is not linearizable against concurrent add() calls – it adds up whatever each cell currently holds, with no synchronization freezing the whole structure first, so a sum() that races with increments can return a value that was never the counter’s value at any single instant. For a request counter or a metric, that’s fine. For anything a decision actually depends on being exact – an inventory count gating “can I sell one more of these” – it is not, and AtomicLong or a lock is the correct choice regardless of throughput.

Going deeper on this section

The cost of writing your own CAS loop

VarHandleCounter does, by hand, exactly what AtomicLong.incrementAndGet() does internally – read, compute, compareAndSet, retry on failure. The benchmark above shows them performing differently under contention anyway: VarHandle‘s hand-written loop holds ~30,000-32,000 ops/ms from 2 threads on, against AtomicLong‘s ~43,000-58,000 – the same operation, a real and repeatable gap.
Reported honestly, not smoothed over. This repository does not have a definitive explanation for the exact size of that gap. The likely cause is that AtomicLong.incrementAndGet() is a method HotSpot recognizes and compiles as an intrinsic, while the hand-written getVolatile + compareAndSet loop in VarHandleCounter, while it issues the same underlying CAS instruction, is ordinary bytecode the JIT has to compile on its own merits. Take the direction seriously – hand-rolling the loop yourself costs something measurable here – and take the exact multiplier as specific to this sandbox rather than a general law.

Going deeper on this section

The ABA problem: when compareAndSet lies to you

compareAndSet checks one thing: is the value still what I last saw? It cannot tell the difference between “nothing changed since I looked” and “it changed to something else and then changed back to exactly what I saw before” – and the second case is not a hypothetical. TreiberStack.java is a textbook lock-free stack (Treiber, 1986) built entirely on one AtomicReference.compareAndSet, and it is exactly where this bites:
1. Thread 1 reads: oldTop=A, newTop=B (stack: A, B, C) – pauses before CAS A B C 2. Main thread pops A, pops B (both real, legitimate pops) C 3. Main thread pushes the SAME “A” object back (pooled/reused node) A C top is “A” again – same object, different structure underneath 4. Thread 1 resumes: CAS(oldTop=A, newTop=B) – top IS A -> succeeds Stack becomes [B, C] – B is back, though it was already removed
$ java -cp target/classes com.ankurm.atomics.AbaProblemDemo

=== Part 1: a real ABA race against TreiberStack ===
Initial stack (top first): [A, B, C]
Main thread popped, legitimately: "A", then "B"
Stack after those two real pops: [C]
Main thread pushed the SAME "A" node object back: [A, C]
Thread 1's stale CAS result: CAS succeeded, pop() would have returned "A"
Stack contents after Thread 1's stale CAS: [B, C]
Output: 02-aba-problem-demo.txt, full commentary included. This isn’t a described race, it’s an actually-run one, forced deterministically with a latch instead of hoped for from real timing: AbaProblemDemo.java pauses Thread 1 right after it reads the top of the stack and right before its CAS, lets the main thread pop twice and push the exact same node object back, then lets Thread 1 finish. "B" reappears in the stack even though nobody pushed it – Thread 1’s CAS matched on reference identity alone and installed a newTop computed from a read that happened before two pops and a push it never saw. "A" also gets handed out twice: once to the main thread’s legitimate pop, once to Thread 1’s stale one – in the object-pool bug that opened this article, that’s the same pooled connection handed to two callers.

Going deeper on this section

Detecting the stale CAS with a stamp

AtomicStampedReference pairs every reference with an integer stamp that the writer bumps on every successful update. A compareAndSet against it has to match both the reference and the stamp it captured – so a reference that went out and came back, like "A" above, no longer matches, even though the reference alone looks identical:
$ java -cp target/classes com.ankurm.atomics.AbaProblemDemo

=== Part 2: AtomicStampedReference detects the same shape of race ===
Reader captured: ref="A", stamp=0
After a concurrent A -> B -> A round trip: ref="A", stamp=2 (reference is back to "A", but the stamp moved on)
A plain AtomicReference.compareAndSet("A", "Z") would see reference == "A" and succeed: true
AtomicStampedReference.compareAndSet("A", "Z", 0, 1) actually succeeded: false (false is correct - the stamp proves a change happened in between, even though the reference alone looks unchanged)
Output: 02-aba-problem-demo.txt. Note what this actually buys: not prevention – the reference genuinely did go "A" -> "B" -> "A", nothing stops that – but detection. The stale compareAndSet fails instead of silently succeeding, which turns a corrupted state into a failed CAS that retries against current, correct data. A stack rewritten on AtomicStampedReference<Node<T>> instead of a plain AtomicReference<Node<T>> would have failed Thread 1’s stale CAS above and forced it to re-read – no corruption, at the cost of one extra integer per reference and a slightly larger API to get right.
Going deeper: do you actually need to worry about this?

Rarely, if you’re not hand-writing lock-free data structures. Every class in java.util.concurrent that could be exposed to ABA – ConcurrentLinkedQueue, ConcurrentLinkedDeque, the Executor framework’s internal queues – was written by people who accounted for it, usually by never reusing node objects the way a pooled allocator does, which is what actually makes ABA possible in the first place: TreiberStack‘s public push/pop in this repo always allocates a fresh node and is not vulnerable at all; the demo above only reproduces the bug because it deliberately reuses a node the way a freelist or object pool would. The situations where this is genuinely your problem: you’re building a lock-free structure yourself with any kind of node pooling or reuse, or you’re using a low-level library that does (some off-heap or native-memory allocators reuse addresses, which is the same shape of bug at the pointer level instead of the object level).

Going deeper on this section

VarHandle’s memory-ordering access modes

Every counter above that isn’t lock-based ultimately reads and writes through some ordering guarantee narrower or equal to volatile. VarHandle exposes that whole spectrum directly, as four named families instead of one keyword:
FIELD.set(demo, 1);                         // plain - no ordering guarantee at all
FIELD.setOpaque(demo, 2);                   // opaque - no reordering with itself, no happens-before with anything else
FIELD.setRelease(demo, 3);                  // release - one-directional happens-before to a later getAcquire
FIELD.setVolatile(demo, 4);                 // volatile - full two-way happens-before, same as a volatile field
Source: VarHandleAccessModesDemo.java. This demo proves the API surface works and returns what its own contract says it returns for a single-threaded run; it cannot and does not prove the ordering guarantees themselves the way the jmm module’s jcstress campaign proved the JLS happens-before rules – that would need the same kind of large concurrent sampling, not a coordination primitive like a latch, so the ordering claims here are attributed to the official VarHandle class documentation rather than to a run captured in this repo:
$ java -cp target/classes com.ankurm.atomics.VarHandleAccessModesDemo

plain set(1) / get() -> 1
setOpaque(2) / getOpaque() -> 2
setRelease(3) / getAcquire() -> 3
setVolatile(4) / getVolatile() -> 4
compareAndSet(4, 5) succeeded -> true, second compareAndSet(4, 6) succeeded -> false (expected false - value is 5, not 4, by the second call)
Output: 03-varhandle-access-modes.txt. The practical reason to know these exist: AtomicLong, AtomicReference and friends are themselves built on VarHandle underneath, in modern JDKs – reaching for the access modes directly only matters when you’re building your own primitive and a full volatile fence on every access is measurably more than the problem needs, which is a genuinely narrow case.

Going deeper on this section

When do atomics actually beat a lock? LongAdder, when the workload is a sum or count under real contention and you don’t need to read it and act on an exact instantaneous value. AtomicLong or AtomicReference, for a single field with simple update logic, where reads must be exact. Not a hand-written VarHandle CAS loop, unless you’re building a library primitive and have measured that you need it – this article’s own benchmark shows it costing something over the built-in atomics for the identical operation. And not a lock-free data structure you wrote yourself with node reuse anywhere in it, unless you’ve specifically reasoned about ABA – the bug that opened this article took a week to reproduce for a reason.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.