diff --git a/README.md b/README.md index 331f377..3118db4 100644 --- a/README.md +++ b/README.md @@ -7,6 +7,7 @@ article; each module's own README has that article's version table, quickstart, |---|---| | [`jmm`](jmm/) | The Java Memory Model Explained: volatile, happens-before, and Why Your Double-Checked Lock Failed | | [`locks`](locks/) | synchronized vs ReentrantLock vs StampedLock: Benchmarks and a Decision Table | +| [`atomics`](atomics/) | Java Atomics and VarHandle: CAS, LongAdder, and When Atomics Beat Locks | ## License diff --git a/atomics/README.md b/atomics/README.md new file mode 100644 index 0000000..b2f5363 --- /dev/null +++ b/atomics/README.md @@ -0,0 +1,78 @@ +# atomics + +Companion code for the ankurm.com post *"Java Atomics and VarHandle: CAS, LongAdder, and When +Atomics Beat Locks."* Third module in `java-core-examples`, the Java-core / concurrency series. + +## Versions this was built and tested against + +| Component | Version | Notes | +|---|---|---| +| JDK | 25.0.4.1+1 (Temurin, LTS) | | +| JMH | 1.37 | | +| JUnit Jupiter | 5.11.0 | Correctness tests only. | +| Maven | 3.9.11 | | +| Hardware | 2 vCPU x86-64 VM | Same sandbox as `jmm` and `locks` - see the honesty note below. | + +## Quickstart + +```bash +export JAVA_HOME=/path/to/jdk-25 +mvn package +java -jar target/benchmarks.jar IncrementBenchmark -t 8 +java -cp target/classes com.ankurm.atomics.AbaProblemDemo +java -cp target/classes com.ankurm.atomics.VarHandleAccessModesDemo +``` + +`scripts/run-all.sh` regenerates every file in `output/`. `scripts/run.sh [args]` +runs one benchmark ad hoc. + +## What's in here + +| File | What it shows | +|---|---| +| `src/main/java/.../Counter.java` | The interface all five counters implement. | +| `.../SynchronizedCounter.java`, `.../ReentrantLockCounter.java` | The two lock-based baselines from the `locks` module, reused here for a direct comparison. | +| `.../AtomicLongCounter.java` | `AtomicLong.incrementAndGet()` - one contended CAS location. | +| `.../LongAdderCounter.java` | `LongAdder` - writes spread across striped cells, `sum()` on read. | +| `.../VarHandleCounter.java` | The same CAS loop `AtomicLong` does internally, written by hand against a plain field via `VarHandle`. | +| `.../TreiberStack.java` | A textbook lock-free stack - and where the ABA problem actually lives. | +| `.../AbaProblemDemo.java` | A real, two-thread, deterministically-forced ABA race against `TreiberStack`, plus the `AtomicStampedReference` fix. | +| `.../VarHandleAccessModesDemo.java` | `VarHandle`'s four access-mode families (plain, opaque, acquire/release, volatile) plus `compareAndSet`. | +| `src/test/java/.../CounterCorrectnessTest.java` | Lost-update sanity checks - does NOT prove throughput, ABA, or ordering claims, see its Javadoc. | +| `output/01` | `IncrementBenchmark` swept across 1, 2, 4, 8, 16, 32, 64 threads. | +| `output/02` | The ABA demo's actual output - a real corruption, then the stamped fix. | +| `output/03` | The VarHandle access-modes demo's output. | +| `output/04` | JUnit correctness run. | + +## Reading the numbers honestly (2-vCPU sandbox) + +**`LongAdder` wins under any real contention** (`output/01`): from 2 threads onward it holds +~160,000-177,000 ops/ms, 3-4x every other counter, exactly matching its design - writes land on +one of several striped cells instead of fighting over one location, so contention on the counter +itself mostly disappears. At 1 thread it's actually the slowest of the five (no contention to +amortize the cell-array bookkeeping against), which is the one case its own Javadoc explicitly +says to expect. + +**`AtomicLong` is fastest at 1 thread, then drops hard and plateaus** (`output/01`): ~152,000 +ops/ms uncontended, down to ~43,000-58,000 from 2 threads on - every thread past the first is +now genuinely fighting over one CAS location, and on this 2-core box that settles into a stable +but much lower plateau rather than degrading further. + +**`VarHandle`'s hand-written CAS loop is consistently behind `AtomicLong`'s under contention** - +roughly 30,000-32,000 ops/ms against `AtomicLong`'s 43,000-58,000 from 2 threads onward, despite +both doing the same fundamental operation (read, compute, `compareAndSet`, retry on failure). +This repository does not have a definitive answer for the exact gap; the likely explanation is +that `AtomicLong.incrementAndGet()` is a JIT/JVM intrinsic HotSpot recognizes and compiles +specially, while the hand-written `getVolatile` + `compareAndSet` loop in `VarHandleCounter`, +while using the same underlying CAS instruction, doesn't get quite the same treatment. Take this +as "hand-rolling the loop yourself costs something measurable here," not as a precise multiplier +to expect elsewhere. + +**`synchronized` again lags well behind at 2+ threads** - the same pattern documented in the +`locks` module's README, reproduced here with a completely different set of counters, which is +some evidence it's a property of this sandbox's monitor-contention path rather than a fluke of +one benchmark file. + +## License + +MIT - see the [repo-wide LICENSE](../LICENSE). diff --git a/atomics/output/01-increment-benchmark-sweep.txt b/atomics/output/01-increment-benchmark-sweep.txt new file mode 100644 index 0000000..b0f0480 --- /dev/null +++ b/atomics/output/01-increment-benchmark-sweep.txt @@ -0,0 +1,59 @@ +$ java -jar target/benchmarks.jar IncrementBenchmark -t -rf text +(one JVM invocation per thread count; -t sets the JMH thread count for that run) + +=== threads=1 === +Benchmark Mode Cnt Score Error Units +IncrementBenchmark.atomicLong thrpt 5 152210.819 ± 13994.910 ops/ms +IncrementBenchmark.longAdder thrpt 5 85474.018 ± 4259.652 ops/ms +IncrementBenchmark.reentrantLock thrpt 5 52749.683 ± 9224.450 ops/ms +IncrementBenchmark.synchronized_ thrpt 5 41962.389 ± 5710.162 ops/ms +IncrementBenchmark.varHandleCas thrpt 5 92813.051 ± 18134.407 ops/ms + +=== threads=2 === +Benchmark Mode Cnt Score Error Units +IncrementBenchmark.atomicLong thrpt 5 55290.174 ± 5756.818 ops/ms +IncrementBenchmark.longAdder thrpt 5 175504.101 ± 19226.430 ops/ms +IncrementBenchmark.reentrantLock thrpt 5 12132.831 ± 3380.323 ops/ms +IncrementBenchmark.synchronized_ thrpt 5 7781.210 ± 5925.389 ops/ms +IncrementBenchmark.varHandleCas thrpt 5 31955.213 ± 7578.028 ops/ms + +=== threads=4 === +Benchmark Mode Cnt Score Error Units +IncrementBenchmark.atomicLong thrpt 5 58192.557 ± 1874.113 ops/ms +IncrementBenchmark.longAdder thrpt 5 176176.145 ± 10500.741 ops/ms +IncrementBenchmark.reentrantLock thrpt 5 49474.377 ± 6131.270 ops/ms +IncrementBenchmark.synchronized_ thrpt 5 8198.887 ± 18413.808 ops/ms +IncrementBenchmark.varHandleCas thrpt 5 32807.244 ± 14075.306 ops/ms + +=== threads=8 === +Benchmark Mode Cnt Score Error Units +IncrementBenchmark.atomicLong thrpt 5 58036.390 ± 10538.772 ops/ms +IncrementBenchmark.longAdder thrpt 5 176823.593 ± 19520.176 ops/ms +IncrementBenchmark.reentrantLock thrpt 5 49528.714 ± 9256.971 ops/ms +IncrementBenchmark.synchronized_ thrpt 5 8933.358 ± 7612.678 ops/ms +IncrementBenchmark.varHandleCas thrpt 5 31671.036 ± 9740.912 ops/ms + +=== threads=16 === +Benchmark Mode Cnt Score Error Units +IncrementBenchmark.atomicLong thrpt 5 55587.800 ± 3218.928 ops/ms +IncrementBenchmark.longAdder thrpt 5 171534.976 ± 37921.512 ops/ms +IncrementBenchmark.reentrantLock thrpt 5 48603.146 ± 6585.079 ops/ms +IncrementBenchmark.synchronized_ thrpt 5 8839.293 ± 3658.010 ops/ms +IncrementBenchmark.varHandleCas thrpt 5 31809.007 ± 7320.695 ops/ms + +=== threads=32 === +Benchmark Mode Cnt Score Error Units +IncrementBenchmark.atomicLong thrpt 5 56397.645 ± 6961.852 ops/ms +IncrementBenchmark.longAdder thrpt 5 162883.357 ± 14305.826 ops/ms +IncrementBenchmark.reentrantLock thrpt 5 43556.152 ± 10557.939 ops/ms +IncrementBenchmark.synchronized_ thrpt 5 9159.280 ± 8780.101 ops/ms +IncrementBenchmark.varHandleCas thrpt 5 32833.760 ± 2661.430 ops/ms + +=== threads=64 === +Benchmark Mode Cnt Score Error Units +IncrementBenchmark.atomicLong thrpt 5 43448.887 ± 5607.031 ops/ms +IncrementBenchmark.longAdder thrpt 5 166468.689 ± 22263.886 ops/ms +IncrementBenchmark.reentrantLock thrpt 5 44277.427 ± 6928.188 ops/ms +IncrementBenchmark.synchronized_ thrpt 5 10784.283 ± 3760.196 ops/ms +IncrementBenchmark.varHandleCas thrpt 5 29751.529 ± 1620.459 ops/ms + diff --git a/atomics/output/02-aba-problem-demo.txt b/atomics/output/02-aba-problem-demo.txt new file mode 100644 index 0000000..54f9352 --- /dev/null +++ b/atomics/output/02-aba-problem-demo.txt @@ -0,0 +1,21 @@ +$ java -cp target/classes com.ankurm.atomics.AbaProblemDemo + +=== Part 1: a real ABA race against TreiberStack === +Initial stack (top first): [A, B, C] +Main thread popped, legitimately: "A", then "B" +Stack after those two real pops: [C] +Main thread pushed the SAME "A" node object back: [A, C] +Thread 1's stale CAS result: CAS succeeded, pop() would have returned "A" +Stack contents after Thread 1's stale CAS: [B, C] +"B" is back in the stack even though the main thread already popped it and +nobody ever pushed it again - Thread 1's CAS matched on reference identity +alone (top was "A" both times it looked) and blindly installed a newTop +("B") that was computed from a read that happened before two pops and a +push it never saw. "A" was also just handed out twice: once to the main +thread's first pop(), once to Thread 1's stale one. + +=== Part 2: AtomicStampedReference detects the same shape of race === +Reader captured: ref="A", stamp=0 +After a concurrent A -> B -> A round trip: ref="A", stamp=2 (reference is back to "A", but the stamp moved on) +A plain AtomicReference.compareAndSet("A", "Z") would see reference == "A" and succeed: true +AtomicStampedReference.compareAndSet("A", "Z", 0, 1) actually succeeded: false (false is correct - the stamp proves a change happened in between, even though the reference alone looks unchanged) diff --git a/atomics/output/03-varhandle-access-modes.txt b/atomics/output/03-varhandle-access-modes.txt new file mode 100644 index 0000000..4e9c9d4 --- /dev/null +++ b/atomics/output/03-varhandle-access-modes.txt @@ -0,0 +1,7 @@ +$ java -cp target/classes com.ankurm.atomics.VarHandleAccessModesDemo + +plain set(1) / get() -> 1 +setOpaque(2) / getOpaque() -> 2 +setRelease(3) / getAcquire() -> 3 +setVolatile(4) / getVolatile() -> 4 +compareAndSet(4, 5) succeeded -> true, second compareAndSet(4, 6) succeeded -> false (expected false - value is 5, not 4, by the second call) diff --git a/atomics/output/04-counter-correctness.txt b/atomics/output/04-counter-correctness.txt new file mode 100644 index 0000000..a945dd4 --- /dev/null +++ b/atomics/output/04-counter-correctness.txt @@ -0,0 +1,4 @@ +------------------------------------------------------------------------------- +Test set: com.ankurm.atomics.CounterCorrectnessTest +------------------------------------------------------------------------------- +Tests run: 5, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 0.221 s -- in com.ankurm.atomics.CounterCorrectnessTest diff --git a/atomics/pom.xml b/atomics/pom.xml new file mode 100644 index 0000000..2b8a7e0 --- /dev/null +++ b/atomics/pom.xml @@ -0,0 +1,83 @@ + + + 4.0.0 + + + com.ankurm + java-core-examples + 1.0 + + + atomics + atomics + Java atomics and VarHandle: a counter written five ways, JMH throughput 1-64 threads, the ABA problem reproduced and fixed, and VarHandle's memory-ordering access modes. + + + 1.37 + + + + + org.openjdk.jmh + jmh-core + ${jmh.version} + + + org.openjdk.jmh + jmh-generator-annprocess + ${jmh.version} + + + org.junit.jupiter + junit-jupiter + 5.11.0 + test + + + + + benchmarks + + + org.apache.maven.plugins + maven-compiler-plugin + 3.13.0 + + 25 + + + org.openjdk.jmh + jmh-generator-annprocess + ${jmh.version} + + + + + + org.apache.maven.plugins + maven-surefire-plugin + 3.2.5 + + + org.apache.maven.plugins + maven-shade-plugin + 3.5.1 + + + package + shade + + + + org.openjdk.jmh.Main + + + + + + + + + diff --git a/atomics/scripts/run-all.sh b/atomics/scripts/run-all.sh new file mode 100755 index 0000000..4a0ff1a --- /dev/null +++ b/atomics/scripts/run-all.sh @@ -0,0 +1,36 @@ +#!/usr/bin/env bash +# Regenerates every file in output/ using the same commands used to produce the ones +# committed here. Requires JDK 25 on PATH/JAVA_HOME. +set -euo pipefail +cd "$(dirname "$0")/.." + +mvn -q package + +echo "--- 01: IncrementBenchmark thread-count sweep (takes a few minutes) ---" +{ + echo '$ java -jar target/benchmarks.jar IncrementBenchmark -t -rf text' + for t in 1 2 4 8 16 32 64; do + echo "=== threads=$t ===" + java -jar target/benchmarks.jar "IncrementBenchmark" -t "$t" -rf text -rff /tmp/jmh-atomics-$t.txt + cat "/tmp/jmh-atomics-$t.txt" + echo "" + done +} > output/01-increment-benchmark-sweep.txt + +echo "--- 02: ABA problem demo ---" +{ + echo '$ java -cp target/classes com.ankurm.atomics.AbaProblemDemo' + echo "" +} > output/02-aba-problem-demo.txt +java -cp target/classes com.ankurm.atomics.AbaProblemDemo >> output/02-aba-problem-demo.txt 2>&1 + +echo "--- 03: VarHandle access modes demo ---" +{ + echo '$ java -cp target/classes com.ankurm.atomics.VarHandleAccessModesDemo' + echo "" +} > output/03-varhandle-access-modes.txt +java -cp target/classes com.ankurm.atomics.VarHandleAccessModesDemo >> output/03-varhandle-access-modes.txt 2>&1 + +echo "--- 04: correctness tests ---" +mvn -q test +cp target/surefire-reports/com.ankurm.atomics.CounterCorrectnessTest.txt output/04-counter-correctness.txt diff --git a/atomics/scripts/run.sh b/atomics/scripts/run.sh new file mode 100755 index 0000000..5d21622 --- /dev/null +++ b/atomics/scripts/run.sh @@ -0,0 +1,13 @@ +#!/usr/bin/env bash +# Quick ad hoc JMH run against one benchmark class. Example: +# ./scripts/run.sh IncrementBenchmark -t 8 +set -euo pipefail +cd "$(dirname "$0")/.." + +BENCH="${1:?Usage: run.sh [extra jmh args]}" +shift || true + +JAR=target/benchmarks.jar +[ -f "$JAR" ] || { echo "Run 'mvn package' first."; exit 1; } + +java -jar "$JAR" "$BENCH" "$@" diff --git a/atomics/src/main/java/com/ankurm/atomics/AbaProblemDemo.java b/atomics/src/main/java/com/ankurm/atomics/AbaProblemDemo.java new file mode 100644 index 0000000..6fa4200 --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/AbaProblemDemo.java @@ -0,0 +1,104 @@ +package com.ankurm.atomics; + +import java.util.concurrent.CountDownLatch; +import java.util.concurrent.atomic.AtomicStampedReference; + +/** + * Two parts. Part 1 forces a real ABA race against {@link TreiberStack} with + * two threads and a latch, so the corruption below is an actual observed + * race outcome, not a described one. Part 2 shows the minimal mechanism + * {@link AtomicStampedReference} uses to detect - not prevent, detect - + * exactly that race. + */ +public final class AbaProblemDemo { + + public static void main(String[] args) throws InterruptedException { + part1TreiberStackAba(); + System.out.println(); + part2StampedReferenceDetectsIt(); + } + + private static void part1TreiberStackAba() throws InterruptedException { + System.out.println("=== Part 1: a real ABA race against TreiberStack ==="); + TreiberStack stack = new TreiberStack<>(); + stack.push("C"); + stack.push("B"); + stack.push("A"); + System.out.println("Initial stack (top first): " + stack.contentsSnapshot()); + + CountDownLatch t1HasRead = new CountDownLatch(1); + CountDownLatch mainHasInterfered = new CountDownLatch(1); + String[] t1Result = new String[1]; + + Thread t1 = new Thread(() -> { + // Read oldTop=A, newTop=B, then pause right before the CAS - exactly + // where a real thread could be preempted for an arbitrarily long time. + TreiberStack.Node[] read = stack.readForPopForDemo(); + TreiberStack.Node oldTop = read[0]; + TreiberStack.Node newTop = read[1]; + t1HasRead.countDown(); + try { + mainHasInterfered.await(); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return; + } + boolean success = stack.finishPopForDemo(oldTop, newTop); + t1Result[0] = success ? ("CAS succeeded, pop() would have returned \"" + oldTop.value + "\"") + : "CAS failed (this is what we WANT to see - it did not happen here)"; + }, "Thread-1-stale-popper"); + t1.start(); + + t1HasRead.await(); // Thread 1 now holds oldTop=A, newTop=B, has not CAS'd yet. + + // Main thread interferes: legitimately pop A, then B (both real pop() calls), + // then push the SAME "A" node object back - simulating a pooled allocator + // that reuses freed nodes instead of always allocating fresh ones. + TreiberStack.Node[] read = stack.readForPopForDemo(); + TreiberStack.Node nodeA = read[0]; + String popped1 = stack.pop(); + String popped2 = stack.pop(); + System.out.println("Main thread popped, legitimately: \"" + popped1 + "\", then \"" + popped2 + "\""); + System.out.println("Stack after those two real pops: " + stack.contentsSnapshot()); + stack.pushSameNodeForDemo(nodeA); // same object identity as Thread 1's oldTop + System.out.println("Main thread pushed the SAME \"A\" node object back: " + stack.contentsSnapshot()); + + mainHasInterfered.countDown(); + t1.join(); + + System.out.println("Thread 1's stale CAS result: " + t1Result[0]); + System.out.println("Stack contents after Thread 1's stale CAS: " + stack.contentsSnapshot()); + System.out.println("\"B\" is back in the stack even though the main thread already popped it and"); + System.out.println("nobody ever pushed it again - Thread 1's CAS matched on reference identity"); + System.out.println("alone (top was \"A\" both times it looked) and blindly installed a newTop"); + System.out.println("(\"B\") that was computed from a read that happened before two pops and a"); + System.out.println("push it never saw. \"A\" was also just handed out twice: once to the main"); + System.out.println("thread's first pop(), once to Thread 1's stale one."); + } + + private static void part2StampedReferenceDetectsIt() { + System.out.println("=== Part 2: AtomicStampedReference detects the same shape of race ==="); + AtomicStampedReference ref = new AtomicStampedReference<>("A", 0); + + int[] stampHolder = new int[1]; + String staleRef = ref.get(stampHolder); + int staleStamp = stampHolder[0]; + System.out.println("Reader captured: ref=\"" + staleRef + "\", stamp=" + staleStamp); + + // Simulate the same A -> B -> A round trip, each transition bumping the stamp - + // exactly what a real concurrent writer would do on every successful update. + ref.set("B", staleStamp + 1); + ref.set("A", staleStamp + 2); + System.out.println("After a concurrent A -> B -> A round trip: ref=\"" + ref.getReference() + + "\", stamp=" + ref.getStamp() + " (reference is back to \"A\", but the stamp moved on)"); + + boolean plainWouldSucceed = staleRef.equals(ref.getReference()); // what a plain == / equals CAS would see + boolean stampedSucceeds = ref.compareAndSet(staleRef, "Z", staleStamp, staleStamp + 1); + + System.out.println("A plain AtomicReference.compareAndSet(\"A\", \"Z\") would see reference == \"A\" and " + + "succeed: " + plainWouldSucceed); + System.out.println("AtomicStampedReference.compareAndSet(\"A\", \"Z\", " + staleStamp + ", " + (staleStamp + 1) + + ") actually succeeded: " + stampedSucceeds + " (false is correct - the stamp proves a change " + + "happened in between, even though the reference alone looks unchanged)"); + } +} diff --git a/atomics/src/main/java/com/ankurm/atomics/AtomicLongCounter.java b/atomics/src/main/java/com/ankurm/atomics/AtomicLongCounter.java new file mode 100644 index 0000000..5acde69 --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/AtomicLongCounter.java @@ -0,0 +1,23 @@ +package com.ankurm.atomics; + +import java.util.concurrent.atomic.AtomicLong; + +/** + * {@link AtomicLong#incrementAndGet()} - a hardware compare-and-swap loop under + * the hood, retried until it succeeds. No lock, no parking, but every thread + * that loses a CAS race spins and retries against the same single contended + * memory location. + */ +public final class AtomicLongCounter implements Counter { + private final AtomicLong count = new AtomicLong(); + + @Override + public void increment() { + count.incrementAndGet(); + } + + @Override + public long get() { + return count.get(); + } +} diff --git a/atomics/src/main/java/com/ankurm/atomics/Counter.java b/atomics/src/main/java/com/ankurm/atomics/Counter.java new file mode 100644 index 0000000..1c19d95 --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/Counter.java @@ -0,0 +1,7 @@ +package com.ankurm.atomics; + +/** The same shared counter, protected five different ways. */ +public interface Counter { + void increment(); + long get(); +} diff --git a/atomics/src/main/java/com/ankurm/atomics/IncrementBenchmark.java b/atomics/src/main/java/com/ankurm/atomics/IncrementBenchmark.java new file mode 100644 index 0000000..db31c09 --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/IncrementBenchmark.java @@ -0,0 +1,50 @@ +package com.ankurm.atomics; + +import org.openjdk.jmh.annotations.*; + +import java.util.concurrent.TimeUnit; + +/** + * All five counters, same write-only workload, run at a fixed thread count + * per JVM invocation via {@code -t N}; {@code scripts/run-all.sh} sweeps 1, + * 2, 4, 8, 16, 32 and 64 threads. + */ +@BenchmarkMode(Mode.Throughput) +@OutputTimeUnit(TimeUnit.MILLISECONDS) +@State(Scope.Benchmark) +@Warmup(iterations = 3, time = 1, timeUnit = TimeUnit.SECONDS) +@Measurement(iterations = 5, time = 1, timeUnit = TimeUnit.SECONDS) +@Fork(1) +public class IncrementBenchmark { + + private final SynchronizedCounter synchronizedCounter = new SynchronizedCounter(); + private final ReentrantLockCounter reentrantLockCounter = new ReentrantLockCounter(); + private final AtomicLongCounter atomicLongCounter = new AtomicLongCounter(); + private final LongAdderCounter longAdderCounter = new LongAdderCounter(); + private final VarHandleCounter varHandleCounter = new VarHandleCounter(); + + @Benchmark + public void synchronized_() { + synchronizedCounter.increment(); + } + + @Benchmark + public void reentrantLock() { + reentrantLockCounter.increment(); + } + + @Benchmark + public void atomicLong() { + atomicLongCounter.increment(); + } + + @Benchmark + public void longAdder() { + longAdderCounter.increment(); + } + + @Benchmark + public void varHandleCas() { + varHandleCounter.increment(); + } +} diff --git a/atomics/src/main/java/com/ankurm/atomics/LongAdderCounter.java b/atomics/src/main/java/com/ankurm/atomics/LongAdderCounter.java new file mode 100644 index 0000000..ce715be --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/LongAdderCounter.java @@ -0,0 +1,26 @@ +package com.ankurm.atomics; + +import java.util.concurrent.atomic.LongAdder; + +/** + * {@link LongAdder} takes the opposite approach to {@link AtomicLongCounter}: + * instead of every thread fighting over one contended CAS location, writes + * are spread across an internal array of per-thread (really, per-probe-hash) + * cells that only get created once contention is actually detected, and + * {@code sum()} adds them all up on read. Writes get cheap; reads get more + * expensive and, crucially, not linearizable with concurrent writes - see + * the README for what that trade-off actually means. + */ +public final class LongAdderCounter implements Counter { + private final LongAdder count = new LongAdder(); + + @Override + public void increment() { + count.increment(); + } + + @Override + public long get() { + return count.sum(); + } +} diff --git a/atomics/src/main/java/com/ankurm/atomics/ReentrantLockCounter.java b/atomics/src/main/java/com/ankurm/atomics/ReentrantLockCounter.java new file mode 100644 index 0000000..5fe6125 --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/ReentrantLockCounter.java @@ -0,0 +1,29 @@ +package com.ankurm.atomics; + +import java.util.concurrent.locks.ReentrantLock; + +/** The explicit-lock baseline, for comparison against the four lock-free strategies. */ +public final class ReentrantLockCounter implements Counter { + private final ReentrantLock lock = new ReentrantLock(); + private long count; + + @Override + public void increment() { + lock.lock(); + try { + count++; + } finally { + lock.unlock(); + } + } + + @Override + public long get() { + lock.lock(); + try { + return count; + } finally { + lock.unlock(); + } + } +} diff --git a/atomics/src/main/java/com/ankurm/atomics/SynchronizedCounter.java b/atomics/src/main/java/com/ankurm/atomics/SynchronizedCounter.java new file mode 100644 index 0000000..a5f840d --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/SynchronizedCounter.java @@ -0,0 +1,16 @@ +package com.ankurm.atomics; + +/** Baseline: the same lock-based approach benchmarked in the {@code locks} module. */ +public final class SynchronizedCounter implements Counter { + private long count; + + @Override + public synchronized void increment() { + count++; + } + + @Override + public synchronized long get() { + return count; + } +} diff --git a/atomics/src/main/java/com/ankurm/atomics/TreiberStack.java b/atomics/src/main/java/com/ankurm/atomics/TreiberStack.java new file mode 100644 index 0000000..2b5bd4d --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/TreiberStack.java @@ -0,0 +1,107 @@ +package com.ankurm.atomics; + +import java.util.concurrent.atomic.AtomicReference; + +/** + * A classic lock-free stack (Treiber, 1986): push and pop both loop on a + * single {@link AtomicReference#compareAndSet} against the top node. This + * textbook implementation is exactly where the textbook ABA problem lives: + * {@code pop()} reads {@code oldTop} and computes {@code newTop} from it, + * and if another thread pops that same node and later pushes the very same + * node object back on - same reference, different {@code next} underneath + * it by then - the CAS below sees the reference it expects and succeeds, + * even though the structure it is about to install ({@code newTop}, + * computed from the now-stale read) is no longer correct. + *

+ * The package-private {@code *ForDemo} methods exist only so + * {@link AbaProblemDemo} can force that exact interleaving deterministically + * - reusing the identical popped {@link Node} object on the way back in, + * which is what a real freelist or pooled-node allocator does and is the + * only way ABA is reproducible on purpose rather than by chance. The public + * {@code push}/{@code pop} API never reuses nodes and is not affected. + */ +public final class TreiberStack { + + static final class Node { + final T value; + volatile Node next; + + Node(T value, Node next) { + this.value = value; + this.next = next; + } + } + + private final AtomicReference> top = new AtomicReference<>(); + + public void push(T value) { + Node oldTop; + Node newNode = new Node<>(value, null); + do { + oldTop = top.get(); + newNode.next = oldTop; + } while (!top.compareAndSet(oldTop, newNode)); + } + + public T pop() { + Node oldTop; + Node newTop; + do { + oldTop = top.get(); + if (oldTop == null) { + return null; + } + newTop = oldTop.next; + } while (!top.compareAndSet(oldTop, newTop)); + return oldTop.value; + } + + public String contentsSnapshot() { + StringBuilder sb = new StringBuilder("["); + Node n = top.get(); + boolean first = true; + int guard = 0; + while (n != null && guard++ < 20) { + if (!first) sb.append(", "); + sb.append(n.value); + first = false; + n = n.next; + } + sb.append("]"); + return sb.toString(); + } + + // --- demo-only access below: never used by push()/pop() above --- + + Node topNodeForDemo() { + return top.get(); + } + + /** Reads oldTop/newTop exactly like pop() does, but stops before the CAS and hands + * both back so the demo can interleave real pop/push calls from another thread + * in between - reproducing the read-side of the race, not simulating it. */ + Node[] readForPopForDemo() { + @SuppressWarnings("unchecked") + Node[] result = new Node[2]; + result[0] = top.get(); // oldTop + result[1] = result[0] == null ? null : result[0].next; // newTop + return result; + } + + /** Completes the CAS a {@link #readForPopForDemo()} call started - this is the + * exact same compareAndSet pop() itself uses, just split in two so the demo can + * inject interference in between. */ + boolean finishPopForDemo(Node oldTop, Node newTop) { + return top.compareAndSet(oldTop, newTop); + } + + /** Pushes back the SAME node object a previous pop observed, exactly as a pooled + * allocator would - the one operation that makes ABA possible. */ + void pushSameNodeForDemo(Node node) { + Node oldTop; + do { + oldTop = top.get(); + node.next = oldTop; + } while (!top.compareAndSet(oldTop, node)); + } +} diff --git a/atomics/src/main/java/com/ankurm/atomics/VarHandleAccessModesDemo.java b/atomics/src/main/java/com/ankurm/atomics/VarHandleAccessModesDemo.java new file mode 100644 index 0000000..0563c64 --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/VarHandleAccessModesDemo.java @@ -0,0 +1,69 @@ +package com.ankurm.atomics; + +import java.lang.invoke.MethodHandles; +import java.lang.invoke.VarHandle; + +/** + * {@link VarHandle} exposes four families of access mode on the same field, + * each a different point on the plain-to-volatile ordering spectrum defined + * by {@code VarHandle}'s own class documentation. This demo runs all four + * against one field and prints what each call returns - it does not and + * cannot prove the ordering guarantees themselves on a 2-vCPU single-run + * demo (that would need the kind of large concurrent campaign the jmm + * module ran with jcstress, not a coordination primitive), so treat this + * as "the API surface actually compiles and does what its Javadoc says + * about its own return values," with the ordering claims themselves + * attributed to the Javadoc in the README and the post. + */ +public final class VarHandleAccessModesDemo { + + private static final VarHandle FIELD; + + static { + try { + FIELD = MethodHandles.lookup() + .findVarHandle(VarHandleAccessModesDemo.class, "value", int.class); + } catch (ReflectiveOperationException e) { + throw new ExceptionInInitializerError(e); + } + } + + @SuppressWarnings("unused") + private volatile int value; + + public static void main(String[] args) { + VarHandleAccessModesDemo demo = new VarHandleAccessModesDemo(); + + // Plain: no ordering or visibility guarantee at all - same as a normal field + // read/write. Fastest, and the only mode allowed to be reordered/cached freely. + FIELD.set(demo, 1); + System.out.println("plain set(1) / get() -> " + (int) FIELD.get(demo)); + + // Opaque: guarantees the write is eventually visible and reads/writes to THIS + // location are not reordered with each other, but gives no happens-before + // relationship with any OTHER variable - "just don't tear or cache forever." + FIELD.setOpaque(demo, 2); + System.out.println("setOpaque(2) / getOpaque() -> " + (int) FIELD.getOpaque(demo)); + + // Acquire/release: the one-directional half of volatile. setRelease publishes + // everything written before it to a thread that later does a getAcquire on the + // same location - one-way happens-before, cheaper than full volatile on some + // hardware because it doesn't need a full bidirectional fence. + FIELD.setRelease(demo, 3); + System.out.println("setRelease(3) / getAcquire() -> " + (int) FIELD.getAcquire(demo)); + + // Volatile: full happens-before both ways, same guarantee as a `volatile` field + // or synchronized access to it - what every Counter in this module's benchmark + // that isn't plain/opaque/acquire-release actually relies on. + FIELD.setVolatile(demo, 4); + System.out.println("setVolatile(4) / getVolatile() -> " + (int) FIELD.getVolatile(demo)); + + // And the compare-and-swap family every lock-free Counter in this module is + // actually built on: + boolean casSucceeded = FIELD.compareAndSet(demo, 4, 5); + boolean casShouldFail = FIELD.compareAndSet(demo, 4, 6); // 4 is stale now, expect false + System.out.println("compareAndSet(4, 5) succeeded -> " + casSucceeded + + ", second compareAndSet(4, 6) succeeded -> " + casShouldFail + + " (expected false - value is 5, not 4, by the second call)"); + } +} diff --git a/atomics/src/main/java/com/ankurm/atomics/VarHandleCounter.java b/atomics/src/main/java/com/ankurm/atomics/VarHandleCounter.java new file mode 100644 index 0000000..8761257 --- /dev/null +++ b/atomics/src/main/java/com/ankurm/atomics/VarHandleCounter.java @@ -0,0 +1,41 @@ +package com.ankurm.atomics; + +import java.lang.invoke.MethodHandles; +import java.lang.invoke.VarHandle; + +/** + * The same compare-and-swap loop {@link AtomicLongCounter} does internally, + * written out by hand against a plain {@code long} field via {@link VarHandle}. + * This is what {@code AtomicLong} is built on, one layer down - no boxing, + * no extra object, just a field and a handle that knows how to fence and + * CAS against it. + */ +public final class VarHandleCounter implements Counter { + + private static final VarHandle COUNT; + + static { + try { + COUNT = MethodHandles.lookup() + .findVarHandle(VarHandleCounter.class, "count", long.class); + } catch (ReflectiveOperationException e) { + throw new ExceptionInInitializerError(e); + } + } + + @SuppressWarnings("unused") + private volatile long count; + + @Override + public void increment() { + long current; + do { + current = (long) COUNT.getVolatile(this); + } while (!COUNT.compareAndSet(this, current, current + 1)); + } + + @Override + public long get() { + return (long) COUNT.getVolatile(this); + } +} diff --git a/atomics/src/test/java/com/ankurm/atomics/CounterCorrectnessTest.java b/atomics/src/test/java/com/ankurm/atomics/CounterCorrectnessTest.java new file mode 100644 index 0000000..f63c2f3 --- /dev/null +++ b/atomics/src/test/java/com/ankurm/atomics/CounterCorrectnessTest.java @@ -0,0 +1,70 @@ +package com.ankurm.atomics; + +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.Timeout; + +import java.util.concurrent.CountDownLatch; +import java.util.stream.IntStream; + +import static org.junit.jupiter.api.Assertions.assertEquals; + +/** + * Lost-update sanity checks for all five counters - does not test throughput + * or the ABA/memory-ordering claims, only that nothing loses an increment. + */ +class CounterCorrectnessTest { + + private static final int THREADS = 8; + private static final int INCREMENTS_PER_THREAD = 50_000; + + @Test @Timeout(30) + void synchronizedCounterHasNoLostUpdates() throws InterruptedException { + assertNoLostUpdates(new SynchronizedCounter()); + } + + @Test @Timeout(30) + void reentrantLockCounterHasNoLostUpdates() throws InterruptedException { + assertNoLostUpdates(new ReentrantLockCounter()); + } + + @Test @Timeout(30) + void atomicLongCounterHasNoLostUpdates() throws InterruptedException { + assertNoLostUpdates(new AtomicLongCounter()); + } + + @Test @Timeout(30) + void longAdderCounterHasNoLostUpdates() throws InterruptedException { + assertNoLostUpdates(new LongAdderCounter()); + } + + @Test @Timeout(30) + void varHandleCounterHasNoLostUpdates() throws InterruptedException { + assertNoLostUpdates(new VarHandleCounter()); + } + + private void assertNoLostUpdates(Counter counter) throws InterruptedException { + CountDownLatch ready = new CountDownLatch(THREADS); + CountDownLatch start = new CountDownLatch(1); + CountDownLatch done = new CountDownLatch(THREADS); + + IntStream.range(0, THREADS).forEach(i -> new Thread(() -> { + ready.countDown(); + try { + start.await(); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return; + } + for (int j = 0; j < INCREMENTS_PER_THREAD; j++) { + counter.increment(); + } + done.countDown(); + }).start()); + + ready.await(); + start.countDown(); + done.await(); + + assertEquals((long) THREADS * INCREMENTS_PER_THREAD, counter.get()); + } +} diff --git a/pom.xml b/pom.xml index cab5ec3..f4273b0 100644 --- a/pom.xml +++ b/pom.xml @@ -15,6 +15,7 @@ jmm locks + atomics