Skip to main content

The Java Memory Model Explained: volatile, happens-before, and Why Your Double-Checked Lock Failed

Why the classic double-checked-locking singleton breaks without volatile, what the Java Memory Model’s happens-before rules actually guarantee, and a real jcstress campaign (5+ billion samples) that reproduces the unsynchronized publish, the volatile fix, and the final-field safe-publication guarantee side by side.

A production Singleton getter is supposed to return the same object every time. In code review, in unit tests, and in months of production traffic, getInstance() did exactly that — until one day, on one particular box, a caller got back an object whose two fields disagreed with each other: one field held the value the constructor set, the other still held zero. Nothing had thrown. Nothing had logged an error. The object was simply wrong, and it was wrong in a way that a debugger attached after the fact could never catch, because by the time anyone looked, the field had already caught up. This is the double-checked-locking bug, and it is not a race in the sense most people mean the word — no two threads write the same field at the same time, no lock is ever skipped by mistake. The code runs exactly as written. What breaks is a much older assumption: that instructions run in the order you wrote them. This article works through why that assumption is false, what the Java Memory Model promises instead, and reproduces the failure — and its two fixes — with a tool built for exactly this, jcstress. Every claim below traces to a real compile, a real run, or the Java Language Specification; a full jcstress campaign that could not reproduce the bug on this hardware is reported as plainly as the ones that could.
Versions. Tested on JDK 25.0.4.1+1 (Temurin, LTS, GA September 2025). jcstress-core 0.16, the newest release on Maven Central at the time of writing (last published February 2023 — the tool is stable, not abandoned). The companion repository also compiles under JDK 21.0.10 to demonstrate one trap below; the jcstress campaigns themselves ran on 25. All jcstress runs were executed on a 2-vCPU x86-64 virtual machine — that detail matters more than usual in this article, and is addressed directly rather than glossed over.

The bug in its smallest possible form

Strip away the singleton, the lock, and the double check, and what’s left is one thread writing an object reference and another thread reading it, with nothing coordinating the two:
// Thread A (writer)                 // Thread B (reader)
instance = new Payload(1, 1);        Payload p = instance;
                                      if (p != null) {
                                          use(p.a, p.b);
                                      }
This sketch strips away the singleton and the lock on purpose, to isolate just the reordering. The real class this article fixes twice further below is BrokenDclSingleton.java, and Payload — the object being published, with deliberately non-final fields — is Payload.java. The intuition almost everyone starts with is: either B runs before the write (sees null) or after it (sees a fully-built Payload with a=1, b=1). Those are the only two outcomes a single-threaded mental model allows, because in a single thread, a constructor obviously finishes before the reference it returns can be stored anywhere.
Thread A (writer) p.a = 1 p.b = 1 instance = p Thread B (reader) p = instance read p.a, p.b solid: the order the program appears to run in dashed: what the JLS actually permits — a reader can see the reference before it sees every field the constructor wrote
That intuition is wrong, and the gap between it and reality is exactly what the rest of this article is about. Nothing in the Java Language Specification requires the three writes in Thread A to become visible to Thread B in the order they were issued — not the compiler’s reordering, not the CPU’s, not the cache’s. The dashed arrow above is not a bug in some particular JVM; it is a reordering the specification explicitly permits when there is no happens-before edge between the writer and the reader. a and b are ordinary fields, not volatile and not final, so nothing forces a happens-before edge to exist, and the reader is allowed to observe the reference before it observes both fields — producing a=1, b=0, a state Thread A’s constructor never actually produced but Thread B is fully entitled to see.

Going deeper on this section

What “happens-before” actually promises

Java does not promise that operations run in program order across threads. What it promises instead is narrower and more useful: a specific, closed list of situations that create a happens-before edge between one action and another. If two actions are connected by a chain of happens-before edges, the first is guaranteed visible to the second, in full. If they are not connected by any such chain, the JLS considers them unordered, and unordered operations on the same field are a data race — visibility is simply not guaranteed in either direction, regardless of how the code reads. The edges that matter for this article, from JLS §17.4.5:
1. Program order — each action happens-before every later action in the same thread 2. Monitor order — an unlock happens-before every subsequent lock on that same monitor 3. Volatile order — a write to a volatile field happens-before every later read of it 4. Thread start/join — Thread.start() happens-before the new thread’s actions, which happen-before a successful join()
Two properties of this list matter more than the individual rules. First, happens-before is transitive: if A happens-before B and B happens-before C, then A happens-before C, even if A and C never touch the same variable. Second, the list is closed — reading the code and seeing that operations “obviously” run in a sensible order is not a happens-before edge. The JLS committee’s own JSR-133 FAQ is blunt about the case that matters here: the double-checked-locking idiom without volatile is broken because “the writes which initialize the object and the write to the instance field can be reordered … which would have the effect of returning what appears to be a partially constructed object.” Every worked example below is one of these four edges either present or absent.
What this is not. This is not about CPU cache coherence in the sense of one core seeing stale data forever — every mainstream CPU eventually propagates a write to every core. It is about ordering: whether the writes a thread issues become visible to another thread in the order they were issued, or the order they were compiled into, versus some other order. Both the compiler and the CPU are free to reorder independent operations when nothing tells them not to, and “nothing tells them not to” is precisely what a plain field is.

Going deeper on this section

Reproducing the unsynchronized publish with jcstress

Talk of reordering is easy to wave at and hard to pin down, so this is where the article stops describing the bug and runs it. jcstress is OpenJDK’s own tool for exactly this: it runs a pair of racing operations across many forked JVMs and millions of interleavings, and classifies every outcome it observes as expected, merely interesting, or forbidden.
@JCStressTest
@Outcome(id = "1, 1", expect = ACCEPTABLE,
        desc = "Reader saw the fully-constructed Payload.")
@Outcome(id = "0, 0", expect = ACCEPTABLE,
        desc = "Reader ran before the write was published at all.")
@Outcome(id = {"1, 0", "0, 1"}, expect = ACCEPTABLE_INTERESTING,
        desc = "Reordering: the reference was visible before both of its fields were.")
@State
public class PlainPublicationTest {
    Payload instance;

    @Actor
    public void writer() {
        instance = new Payload(1, 1);
    }

    @Actor
    public void reader(II_Result r) {
        Payload p = instance;
        if (p == null) { r.r1 = 0; r.r2 = 0; }
        else { r.r1 = p.a; r.r2 = p.b; }
    }
}
Source: PlainPublicationTest.java. This is literally the unsynchronized publish inside BrokenDclSingleton.getInstance(), isolated from the double-checking logic so jcstress can hammer on just the part that matters. II_Result is jcstress’s own two-int result type — its result classes are named by type code, not by a generic IntResult2, which is worth knowing before you go looking for one in the Javadoc.
$ java -jar target/jcstress.jar -t PlainPublicationTest -jvmArgs "-Xmx256m" -f 2 -iters 15 -time 1000

RUN RESULTS:
  Interesting tests: No matches.
  Failed tests: No matches.

  RESULT        SAMPLES     FREQ       EXPECT  DESCRIPTION
    0, 0  3,511,549,619   66.79%   Acceptable  Reader ran before the write was published at all.
    0, 1              0    0.00%  Interesting  Reordering: the reference was visible before both of its ...
    1, 0              0    0.00%  Interesting  Reordering: the reference was visible before both of its ...
    1, 1  1,745,742,499   33.21%   Acceptable  Reader saw the fully-constructed Payload.
Output: 01-plain-publication.txt. Over roughly 5.26 billion sampled interleavings, the torn outcome never showed up on this hardware.
That “never” is reported honestly, not swept aside. This sandbox is a 2-vCPU x86-64 VM. x86’s TSO memory model does not permit the CPU itself to reorder one store ahead of another store from the same thread — which is half of why the classic “works on x86, fails on ARM” story is true: ARM and POWER permit store-store reordering in hardware, x86 does not. The other half is compiler-level reordering, which is legal on any architecture and is what the specification is actually defending against, but which a particular JIT may or may not trigger for a particular two-field object on a particular run. Neither this repository nor this article claims to have reproduced the failure on ARM silicon — there wasn’t any available to test on. The claim that DCL fails on weak-memory hardware is attributed to the JSR-133 rationale and the historical record (see Further reading), not to a run captured here. What this campaign does prove directly: jcstress classifies the torn outcome as merely Interesting here, not ruled out — contrast that with the two fixes below, where the same outcome is marked Forbidden and a single occurrence would show up as a hard error.

Going deeper on this section

The 2004 fix: one keyword, one happens-before edge

JSR-133 (2004) rewrote the Java Memory Model specifically to fix this idiom, and the fix is one word: mark the reference volatile. A write to a volatile field happens-before every subsequent read of that same field (rule 3 above) — which chains with program order (rule 1) into a complete path from the constructor’s field writes all the way to the reader’s field reads:
p.a = 1 p.b = 1 instance = p (volatile) p = instance (volatile) read p.a, p.b every link solid — transitivity carries visibility all the way through
public final class VolatileDclSingleton {
    private static volatile Payload instance;

    public static Payload getInstance() {
        if (instance == null) {
            synchronized (VolatileDclSingleton.class) {
                if (instance == null) {
                    instance = new Payload(1, 1);
                }
            }
        }
        return instance;
    }
}
Source: VolatileDclSingleton.java. The matching jcstress test, VolatilePublicationTest.java, is PlainPublicationTest with one field made volatile — which changes the torn outcomes’ classification from Interesting to Forbidden, meaning jcstress will fail the run outright if it ever sees one:
$ java -jar target/jcstress.jar -t VolatilePublicationTest -jvmArgs "-Xmx256m" -f 2 -iters 15 -time 1000

RUN RESULTS:
  Interesting tests: No matches.
  Failed tests: No matches.

  RESULT        SAMPLES     FREQ      EXPECT  DESCRIPTION
    0, 0  4,290,607,784   80.26%  Acceptable  Reader ran before the write was published at all.
    0, 1              0    0.00%   Forbidden  Would mean the volatile happens-before edge failed to hold.
    1, 0              0    0.00%   Forbidden  Would mean the volatile happens-before edge failed to hold.
    1, 1  1,055,076,014   19.74%  Acceptable  Reader saw the fully-constructed Payload.
Output: 02-volatile-publication.txt. Zero forbidden outcomes across roughly 5.35 billion samples, exactly as the specification requires — this is the one claim in this article that would be genuinely alarming if it came out any other way, since it would mean a JVM violating its own memory model.
Going deeper: why the second check inside the lock is still needed

If volatile alone fixes the visibility problem, why does the code still check instance == null twice? Visibility and mutual exclusion are two different problems. The outer check is a fast path — most calls, after the first, just read a non-null volatile and return, at the cost of one volatile read. The synchronized block still exists because two threads can both observe null on the outer check before either has taken the lock; without the inner check, both would proceed to construct a Payload, and whichever write happened last would silently win, handing different callers different objects for a supposed singleton. The inner check, taken under the lock, is what guarantees only one Payload is ever constructed. volatile fixes what the other threads see once the object exists; the lock fixes how many objects get built in the first place.

Going deeper on this section

The fix that needs no volatile at all

There is a second, older way out of this, and it doesn’t touch the reference field: make the payload’s own fields final.
public final class FinalPayload {
    public final int a;
    public final int b;

    public FinalPayload(int a, int b) {
        this.a = a;
        this.b = b;
    }
}
Source: FinalPayload.java. JLS §17.5 gives final fields a guarantee that has nothing to do with volatile: if an object is correctly constructed — meaning no reference to it escapes while the constructor is still running — then any thread that later obtains a reference to it, by any means at all, is guaranteed to see the values its final fields were given in the constructor. The reference field carrying that object can be as plain and unsynchronized as you like; the guarantee travels with the object, not with the pointer to it.
$ java -jar target/jcstress.jar -t FinalFieldPublicationTest -jvmArgs "-Xmx256m" -f 2 -iters 15 -time 1000

RUN RESULTS:
  Interesting tests: No matches.
  Failed tests: No matches.

  RESULT        SAMPLES     FREQ      EXPECT  DESCRIPTION
    0, 0  3,697,560,129   66.13%  Acceptable  Reader ran before the write was published at all.
    0, 1              0    0.00%   Forbidden  Would mean the final-field safe-publication guarantee fai...
    1, 0              0    0.00%   Forbidden  Would mean the final-field safe-publication guarantee fai...
    1, 1  1,893,699,349   33.87%  Acceptable  Reader saw the fully-constructed FinalPayload.
Output: 03-final-field-publication.txt. Same shape of test as the two runs above, same zero-forbidden-outcomes result — but note the reference field in FinalFieldPublicationTest.java is a plain, non-volatile field. The guarantee here comes entirely from the payload’s fields being final, not from anything protecting the pointer to it. This is exactly the mechanism the initialization-on-demand holder idiom leans on, and it needs neither volatile nor an explicit lock:
public final class HolderIdiomSingleton {
    private static final class Holder {
        static final Payload INSTANCE = new Payload(1, 1);
    }

    public static Payload getInstance() {
        return Holder.INSTANCE;
    }
}
Source: HolderIdiomSingleton.java. Here the safety comes from a third source entirely — JLS §12.4.2’s class-initialization lock. The JVM guarantees a class’s <clinit> runs exactly once, that every other thread touching the class blocks until it finishes, and that the end of <clinit> happens-before every subsequent use of the class. Holder.INSTANCE does not even need to be final for this particular guarantee — the edge comes from class initialization, not the field modifier — though marking it final documents the intent for free.
Thread A: runs <clinit> Thread B: blocked <clinit> completes A: Holder.INSTANCE B: Holder.INSTANCE the JVM enforces this edge for every class, with no code required from you

Going deeper on this section

Two traps that cost real time building the code above

Both of these came from actually building the companion repository, not from reading about jcstress secondhand, and both are worth knowing before you reach for this tool yourself. jcstress tests have to live in src/main/java, not src/test/java. jcstress ships as a self-executing shaded jar built by maven-shade-plugin, and shade only bundles the main artifact — anything under src/test never makes it into target/jcstress.jar.
The fingerprint of this one. The build succeeds. The jar is produced. Only when you run it does java -jar target/jcstress.jar -t YourTest throw a bare java.lang.NullPointerException from org.openjdk.jcstress.infra.runners.TestList.getTests(), with no mention of your test class anywhere in the stack trace. That specific exception, from that specific class, means the shaded jar has zero tests registered — check that your @JCStressTest classes are under src/main before looking anywhere else.
JDK 25 stopped discovering annotation processors on the plain classpath, silently. jcstress generates its actual test harness (a class named YourTest_jcstress) via an annotation processor at compile time. On JDK 21, javac still finds that processor from the compile classpath with no special configuration — with a warning:
Note: Annotation processing is enabled because one or more processors were found
  on the class path. A future release of javac may disable annotation processing
  unless at least one processor is specified by name (-processor), or a search
  path is specified (--processor-path, --processor-module-path), or annotation
  processing is enabled explicitly (-proc:only, -proc:full).
On JDK 25, that future arrived: the same command produces no warning, no error, and no generated harness class. The build reports success, and mvn package quietly produces a jar with nothing runnable in it. The fix is to declare jcstress-core explicitly as an <annotationProcessorPath> in the compiler plugin configuration, which works on both JDKs:
<annotationProcessorPaths>
  <path>
    <groupId>org.openjdk.jcstress</groupId>
    <artifactId>jcstress-core</artifactId>
    <version>0.16</version>
  </path>
</annotationProcessorPaths>
Output: 04-annotation-processing-jdk21-vs-jdk25.txt, captured directly with javac on both JDKs against this repository’s own source. Source of the fix: jmm/pom.xml.

Going deeper on this section

  • Official reference: JEP 498 and JEP 471, the two JEPs behind the next section’s timeline, which touch annotation processing only tangentially but explain the direction javac is moving in generally

What’s quietly changing underneath all of this in 2026

Two developments this year touch the same territory as this article, and neither is the kind of thing you’d notice unless you were already looking at memory visibility and unsafe publication. The first is that sun.misc.Unsafe‘s memory-access methods — the ones people historically reached for to hand-roll exactly the kind of fences this article discusses — are being phased out on a multi-release schedule that has now reached the runtime-warning stage:
JDK 23@Deprecated(forRemoval)no runtime effect JDK 24runtime warningdefault: warn JDK 25 – 26still warn by defaultdeny available via flag future releasedefault becomes denythrows at call time verified against JEP 471, JEP 498, and this article’s own JDK 25 build (see the previous section)
The recommended replacement for the on-heap methods is exactly the API this article already used the safer, standard-library version of implicitly: VarHandle (JDK 9), which exposes the same plain / opaque / acquire-release / volatile access modes this article’s happens-before rules describe, as a typed, supported API rather than a hidden field-offset trick. MemorySegment (JDK 22’s Foreign Function & Memory API) covers the off-heap methods. The second development is more directly relevant to the two singleton fixes above: JEP 531, Lazy Constants (third preview, targeting JDK 27) proposes a JDK-native LazyConstant-style API that computes a value on first use, exactly once, with the same safe-publication guarantee the holder idiom above gets from class initialization — as a preview feature, not yet something to build on, but the direction both hand-written idioms in this article are heading. That article covers it end to end, including the failure-handling rule that changed between the JDK 26 and 27 previews; this one stops at “it exists and here is why it needs the guarantee this article just proved.”

Going deeper on this section

Should you ever write double-checked locking by hand?

No — reach for the holder idiom, or an AtomicReference/ConcurrentHashMap.computeIfAbsent for anything that isn’t a static singleton. The volatile fix is correct, and it is worth understanding, but every reason to write it by hand today is a reason the holder idiom already covers with less code and no reasoning about memory ordering required from the next person to touch the file. The one place hand-written double-checked locking still turns up honestly is caching libraries and frameworks written before the holder idiom was common knowledge, or that need lazy initialization of something that is not a static field (an instance-level cache, for example, where the holder idiom’s class-loading trick does not apply) — AtomicReference.updateAndGet or ConcurrentHashMap.computeIfAbsent cover that case without a hand-rolled volatile field. If you inherit code with this pattern, the fix is not necessarily a rewrite: confirm the reference field is volatile, and if it is, the code is correct even if it looks dated.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.