Debugging OutOfMemoryError: Heap Dumps, Eclipse MAT, and the Five Leak Patterns in Spring Apps
Five real Spring Boot apps, each driven to a genuine OutOfMemoryError under a constrained heap, each with a real .hprof dump read through Eclipse Memory Analyzer’s headless batch report generator. Static caches, ThreadLocal, unremoved listeners, classloader leaks, and unbounded queues — with the real console transcripts and MAT leak-suspect evidence for every one, plus what the JVM’s own defaults quietly leave out when the crash happens in a container.
Your Spring Boot service has been fine for three weeks. Then, at 2am, it is not: response times climb, then the pod restarts, and the only evidence is a line in the logs that says java.lang.OutOfMemoryError: Java heap space — or worse, no line at all, just Killed and an exit code of 137, because Kubernetes got there first. Nobody changed the code that week. Nothing in the diff looks wrong. And now you have a 1.5 GB .hprof file nobody on the team has opened before.
This is the one bug class where reading the stack trace does not help, because there usually isn’t one that points at the cause — only at the unlucky line that happened to be allocating when the heap finally ran out. Finding the actual leak means reading the heap itself: what is alive, what is holding onto it, and why the garbage collector cannot let it go.
This article drives five real Spring Boot applications to a real OutOfMemoryError under a constrained heap, captures the real .hprof each crash leaves behind, and runs every one of them through Eclipse Memory Analyzer’s (MAT) headless batch report generator for real leak-suspect evidence. Nothing below is a description of what a leak would look like in a dump — every console line and every MAT paragraph quoted here is copied verbatim out of a file a script produced, committed in the companion repository, asmhatre/spring-boot-demo-oom.
Versions. JDK 25.0.4.1+1 (Temurin, LTS) · Spring Boot 4.1.1 (the newest non-milestone GA at write time, verified against repo1.maven.org‘s maven-metadata.xml for spring-boot-starter-parent — its own <release> element currently points at the 4.2.0-M2 milestone, which is not a release) · Eclipse Memory Analyzer 1.17.0 (2026-06-01 RCP build) · Maven 3.9.11.
What “out of memory” actually means
Java’s garbage collector does not go looking for leaks. It only ever asks one question about an object: can I still reach you, starting from a small fixed set of places called GC roots — static fields, local variables currently on a thread’s stack, and a handful of JVM-internal references? If the answer is yes, the object is live, full stop, no matter how obviously useless it is to your program. A Java “memory leak” is never a forgotten free(); it is a reference graph where something, usually by accident, still points at an object a human would call garbage.
A leak, then, is always one of two shapes: either something is stuck directly in a GC root (a static field is itself a root, so anything it still points to is live forever by construction), or an otherwise-unreachable object has one surviving path back to a root that a human did not intend — a stale cache entry, a listener nobody removed, a thread-local slot nobody cleared. Every pattern in this article is a different concrete way of building that one surviving path. Watch for it: the mental model does not change across the five, only the shape of the path does.
Oracle’s Garbage Collector tuning guide covers root sets and generational collection in more depth than this article needs up front.
The smallest leak you can watch happen
static-cache-leak is the simplest of the five modules and the one to start with: a single class, a single field, no threads, no listeners, nothing hiding the mechanism. It is a response cache that somebody added without an eviction policy, which is a bug that has shipped to production under a different class name more times than any of the other four patterns here.
static final Map<String, byte[]> RESPONSE_CACHE = new ConcurrentHashMap<>();
@Bean
ApplicationRunner driveLeak() {
return (ApplicationArguments args) -> {
long payloadSize = 64 * 1024; // 64 KB "cached response" per unique request
long i = 0;
while (true) {
String key = "req-" + Instant.now().toEpochMilli() + "-" + i;
byte[] payload = new byte[(int) payloadSize];
payload[0] = (byte) (i % 128);
RESPONSE_CACHE.put(key, payload);
i++;
}
};
}
The whole bug is three words: static, no eviction, unbounded key space. Every request builds a new key (a timestamp plus a counter, standing in for a request id, a session id, a cache-busting query string — anything with high cardinality), so RESPONSE_CACHE grows by one entry forever. ConcurrentHashMap only makes it thread-safe to leak from several threads at once; it is not the bug. Full file: StaticCacheLeakApplication.java.
Run it under a heap capped at 160 MB, with -XX:+HeapDumpOnOutOfMemoryError so the crash leaves a real dump behind, and this is the complete, unedited transcript — from docs/output/01-oom-console.txt:
$ java -Xmx160m -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=docs/output/heap.hprof -jar target/static-cache-leak.jar
static-cache-leak: entries=1400 approxRetainedMB=87
static-cache-leak: entries=1600 approxRetainedMB=100
static-cache-leak: entries=1800 approxRetainedMB=112
static-cache-leak: entries=2000 approxRetainedMB=125
static-cache-leak: entries=2200 approxRetainedMB=137
java.lang.OutOfMemoryError: Java heap space
Dumping heap to docs/output/heap.hprof ...
Heap dump file created [166360462 bytes in 0.172 secs]
That is a real JVM, really out of heap, and a real 158 MB .hprof on disk 0.172 seconds later. Now the actual question starts: which object is holding the memory? Eclipse MAT’s headless batch mode answers it without ever opening the GUI — ParseHeapDump.sh <dump> org.eclipse.mat.api:suspects parses the dump and writes an HTML leak-suspects report, the same one you would get by right-clicking the dump in the Eclipse MAT UI and choosing Leak Suspects. Here is Problem Suspect 1, trimmed and re-wrapped for width but otherwise verbatim, from docs/output/02-mat-leak-suspects.txt:
Problem Suspect 1 The class com.ankurm.oomdemo.StaticCacheLeakApplication , loaded by
org.springframework.boot.loader.launch.LaunchedClassLoader @ 0xf6067460 , occupies 150,878,464
(96.17%) bytes. The top consumers of its minimum retained heap are byte[] (4,607 instances
totaling 150,730,920), java.util.concurrent.ConcurrentHashMap$Node (2,298 instances totaling
73,536), and java.lang.String (2,309 instances totaling 55,416). The memory is accumulated in
one instance of java.util.concurrent.ConcurrentHashMap$Node[] , loaded by <system class loader>
, which occupies 150,875,504 (96.17%) bytes.
That paragraph is MAT’s dominator tree talking, and it says exactly what the source code says: one class (because a static field is reachable from the class itself, which is a GC root once loaded) retains 96.17% of the live heap, and the instances actually taking the space are 4,607 byte[] objects — the cached payloads — sitting inside a ConcurrentHashMap$Node[], which is the hash table backing RESPONSE_CACHE. Nothing about this required reading a stack trace; it fell out of the dump.
MAT’s headless mode still needs a display.ParseHeapDump.sh runs with no GUI window and produces a static HTML report, but MAT’s SWT runtime initializes a display connection regardless — it fails outright on a headless Linux box without one. Wrap the call in xvfb-run -a and it behaves exactly like the interactive Leak Suspects wizard, just non-interactively. See mat-report.sh for the exact invocation used for all five dumps in this article.
Going deeper: the full MAT Problem Suspect 1 paragraph for static-cache-leak
Problem Suspect 1 The class com.ankurm.oomdemo.StaticCacheLeakApplication , loaded by
org.springframework.boot.loader.launch.LaunchedClassLoader @ 0xf6067460 , occupies 150,878,464
(96.17%) bytes. The top consumers of its minimum retained heap are byte[] (4,607 instances
totaling 150,730,920), java.util.concurrent.ConcurrentHashMap$Node (2,298 instances totaling
73,536), and java.lang.String (2,309 instances totaling 55,416). The memory is accumulated in
one instance of java.util.concurrent.ConcurrentHashMap$Node[] , loaded by <system class loader>
, which occupies 150,875,504 (96.17%) bytes. Thread java.lang.Thread @ 0xf60af750 main has a
local variable or reference to class com.ankurm.oomdemo.StaticCacheLeakApplication @ 0xf6000000
which is on the shortest path to java.util.concurrent.ConcurrentHashMap$Node[4096] @ 0xfba808c8
. The thread java.lang.Thread @ 0xf60af750 main keeps local variables with total size 109,016
(0.07%) bytes. The top consumers of its minimum retained heap are byte[] (4,607 instances
totaling 150,730,920), java.util.concurrent.ConcurrentHashMap$Node (2,298 instances totaling
73,536), and java.lang.String (2,309 instances totaling 55,416). Significant stack frames and
local variables org.springframework.boot.SpringApplication.run(Ljava/lang/Class;[Ljava/lang/Stri
ng;)Lorg/springframework/context/ConfigurableApplicationContext; (SpringApplication.java:1354)
class com.ankurm.oomdemo.StaticCacheLeakApplication @ 0xf6000000 retains 150,878,464 (96.17%)
bytes org.springframework.boot.loader.launch.Launcher.launch(Ljava/lang/ClassLoader;Ljava/lang/S
tring;[Ljava/lang/String;)V (Launcher.java:106) class
com.ankurm.oomdemo.StaticCacheLeakApplication @ 0xf6000000 retains 150,878,464 (96.17%) bytes
The stacktrace of this Thread is available. See stacktrace . See stacktrace with involved local
variables .
Full file: 02-mat-leak-suspects.txt. Regenerate it yourself with run.sh then mat-report.sh.
One intermediate-level fact worth carrying into every section after this one: MAT’s “occupies” figure is a retained size, not a shallow size — it is the total memory that would become collectible if this one object disappeared, computed from the heap’s dominator tree rather than by just adding up field sizes. A ConcurrentHashMap instance itself is a few dozen bytes; its retained size here is 150 MB, because every entry, every node, and every payload is only reachable through it. The next section stays with that idea before moving to pattern two.
Repository: static-cache-leak/ — source, pom.xml, scripts, and both captured transcripts.
MAT’s “Problem Suspect” paragraph is a convenience wrapper around one structure: the heap’s dominator tree. Object A dominates object B if every path from every GC root to B happens to pass through A. The dominator tree is just that relationship made explicit as a tree, and it is why MAT can say a single ConcurrentHashMap “retains” 150 MB even though no single field on it is anywhere near that large — retained size is the sum of everything that becomes unreachable the moment that one node is removed.
In practice that means: start every leak investigation from MAT’s Leak Suspects report (the batch command this article uses throughout, org.eclipse.mat.api:suspects), because it has already walked the dominator tree and sorted by retained size for you. Only drop into the interactive dominator tree view in the full GUI when the suspect report’s top candidate is ambiguous — which did not happen once across these five dumps, all of which had one object retaining between 82% and 96% of the live heap.
xvfb-run -a "$MAT_HOME/ParseHeapDump.sh" "$MODULE_DIR/docs/output/heap.hprof" org.eclipse.mat.api:suspects
That one line, the batch-mode equivalent of opening a dump and choosing Leak Suspects from the menu, is what generated every MAT paragraph in this article — no interactive session involved. Full script, including the result re-extraction: mat-report.sh.
Sort the histogram by retained size, not shallow size. MAT’s plain class histogram defaults to shallow size, which only ranks classes by how many instances exist and how big each one is on its own. A one-line wrapper class with a thousand tiny instances will out-rank the single map actually holding 150 MB under that sort. The Leak Suspects report used throughout this article already sorts by retained size for you; if you ever drop into the histogram view directly, re-sort it before trusting the order.
With the workflow established, the remaining four patterns move faster: each gets the code, the real crash, and the real MAT evidence, but the heap-dump mechanics are the same every time.
Going deeper on the batch tooling: extract-mat-summary.py, the plain Python script that re-extracts the Problem Suspect paragraph from MAT’s generated .zip report — no hand-retyping anywhere in the pipeline.
Pattern two: a ThreadLocal that outlives the request
threadlocal-leak runs a fixed four-thread pool, exactly like a web server’s request-handling pool, and leaks on every single call without ever touching a collection you would think to inspect.
private static void handleOneRequest(long requestId, long payloadSize) {
ThreadLocal<byte[]> requestContext = new ThreadLocal<>();
requestContext.set(new byte[(int) payloadSize]);
requestContext.get()[0] = (byte) requestId; // "use" the context
// Missing on purpose: requestContext.remove();
}
The bug is one missing line, and the mechanism is worth getting exactly right, because it is the part of this pattern people usually get wrong. A Thread has a private ThreadLocalMap, and each entry’s key — the ThreadLocal instance itself — is held by a WeakReference. So once requestContext goes out of scope here, the key really can be collected. But ThreadLocalMap.Entry extends WeakReference<ThreadLocal<?>> while holding the value — the 256 KB byte[] — with an ordinary strong reference. A stale entry (key already null) is only expunged the next time code calls get(), set(), or remove() on that exact slot. Because every request here creates a brand-new ThreadLocal, no request ever revisits an old slot, so nothing ever triggers the cleanup. Each of the four pooled threads accumulates one dead-keyed, live-valued entry per request, forever.
Full file: ThreadLocalLeakApplication.java. Under a 220 MB heap, the leak reached a real OutOfMemoryError thrown on a real worker thread, with a clean stack trace into the leaking call — from docs/output/01-oom-console.txt:
threadlocal-leak: submitted=3017200
threadlocal-leak: submitted=3017600
threadlocal-leak: submitted=3018000
threadlocal-leak: submitted=3018400
threadlocal-leak: submitted=3018800
java.lang.OutOfMemoryError: Java heap space
Dumping heap to docs/output/heap.hprof ...
Heap dump file created [300071460 bytes in 1.908 secs]
Exception in thread "req-worker-1" java.lang.OutOfMemoryError: Java heap space
at com.ankurm.oomdemo.ThreadLocalLeakApplication.handleOneRequest(ThreadLocalLeakApplication.java:61)
at com.ankurm.oomdemo.ThreadLocalLeakApplication.lambda$driveLeak$1(ThreadLocalLeakApplication.java:47)
at com.ankurm.oomdemo.ThreadLocalLeakApplication$$Lambda/0x000000002d25dce8.run(Unknown Source)
at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1090)
at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:614)
at java.base/java.lang.Thread.runWith(Thread.java:1487)
at java.base/java.lang.Thread.run(Thread.java:1474)
Over three million requests landed before the crash — the per-entry cost here is small, which is exactly why this pattern survives in production for weeks before anyone notices. MAT’s dominator tree, from docs/output/02-mat-leak-suspects.txt, does not point at ThreadLocal at all — it points at the pool:
Problem Suspect 1 One instance of
org.springframework.scheduling.concurrent.ThreadPoolTaskExecutor$1 loaded by
org.springframework.boot.loader.launch.LaunchedClassLoader @ 0xf2434790 occupies 172,593,776
(82.46%) bytes. The top consumers of its minimum retained heap are
com.ankurm.oomdemo.ThreadLocalLeakApplication$$Lambda+0x000000006f25dce8 (3,082,020 instances
totaling 98,624,640), java.util.concurrent.LinkedBlockingQueue$Node (3,082,021 instances
totaling 73,968,504), and java.util.HashMap$Node (4 instances totaling 128).
That is the correct diagnosis, and it is a useful lesson on its own: the dominator tree finds the thread (via Spring’s executor wrapper) that the leaking ThreadLocalMap entries are reachable from, not the ThreadLocal class, because the class itself is already gone. When a MAT report fingers a thread pool or an executor class in a leak you did not expect, check for exactly this shape before anything else.
Virtual threads change this specific calculus, not the rule. A virtual thread created per task and never pooled discards its ThreadLocalMap entirely when the task ends, so a ThreadLocal exactly like this one cannot accumulate across requests the way it does on a reused platform thread. That is not a reason to skip remove() — an InheritableThreadLocal captured by a long-lived parent, or a ThreadLocal on a thread from a virtual-thread-unaware pool, leaks exactly the same way. Treat virtual threads as removing one failure mode, not the whole pattern.
ThreadLocal javadoc — the class whose own documentation recommends try/finally with remove().
Pattern three: the listener nobody unregisters
listener-leak is the shape a cache-invalidation bus, a metrics listener, or a Spring ApplicationListener subscriber list takes when its removal path was never written. A long-lived singleton list only ever grows; a session-scoped object registers a lambda with it on “open” and never calls the matching remove on “close”.
static final List<Consumer<String>> LISTENERS = new CopyOnWriteArrayList<>();
private static void openAndCloseOneSession(long sessionId, long payloadSize) {
byte[] sessionData = new byte[(int) payloadSize];
sessionData[0] = (byte) sessionId;
Consumer<String> handler = (String event) -> {
if (sessionData.length > 0 && event != null && event.isEmpty()) {
System.out.print("");
}
};
LISTENERS.add(handler);
// ... session does its work here ...
// Session "closes" -- but nobody ever does LISTENERS.remove(handler).
}
The session object itself, sessionData, can fall completely out of scope the moment this method returns — nothing outside holds a reference to it. It still cannot be collected, because handler is a lambda that closes oversessionData, and LISTENERS holds a strong reference to handler forever. The “session” concept the reader of the code has in their head goes away; the object graph does not agree. Full file: ListenerLeakApplication.java.
listener-leak: driving LISTENERS to OOM via unregistered session listeners
listener-leak: sessions opened=300 live listeners=300
listener-leak: sessions opened=600 live listeners=600
listener-leak: sessions opened=900 live listeners=900
java.lang.OutOfMemoryError: Java heap space
Dumping heap to docs/output/heap.hprof ...
Heap dump file created [156718029 bytes in 0.137 secs]
Exception: java.lang.OutOfMemoryError thrown from the UncaughtExceptionHandler in thread "main"
The fingerprint worth memorizing.live listeners climbing in exact lockstep with sessions opened, 1:1, forever. A listener, subscriber, or callback count that only ever goes up, with no corresponding drop on the equivalent of “session close”, is this bug before you ever open a debugger.
Problem Suspect 1 The class com.ankurm.oomdemo.ListenerLeakApplication , loaded by
org.springframework.boot.loader.launch.LaunchedClassLoader @ 0xf60535d8 , occupies 141,336,208
(95.93%) bytes. The top consumers of its minimum retained heap are byte[] (1,085 instances
totaling 141,313,240), com.ankurm.oomdemo.ListenerLeakApplication$$Lambda+0x000000009b25dce8
(1,079 instances totaling 17,248), and java.lang.Object[] (6 instances totaling 4,504). The
memory is accumulated in one instance of java.lang.Object[] , loaded by <system class loader> ,
which occupies 141,334,440 (95.93%) bytes.
java.lang.Object[] here is CopyOnWriteArrayList‘s backing array — note that it, not a Node type, is what shows up, because CopyOnWriteArrayList is a plain array under a copy-on-write discipline, not a linked structure. Over a thousand lambda instances, each pinning its own session’s payload, all reachable from one array.
Spring’s own event system has the identical failure mode: an ApplicationListener bean registered once lives for the application’s lifetime by design, which is correct for a singleton listener, but a listener built per-request (say, inside a @RequestScope bean that subscribes to a shared ApplicationEventPublisher-backed bus in its constructor) needs an explicit unsubscribe in a @PreDestroy or it reproduces exactly this graph. The framework will not warn you — subscriber lists are supposed to be append-friendly.
Pattern four: the classloader that will not let go
classloader-leak demonstrates the pattern behind the single worst memory leak category in application servers: hot-reloaded code whose old classloader never gets released. The module reloads the same tiny “plugin” class, PluginTask, through a brand-new ClassLoader every iteration — exactly what a plugin system, a rules-engine hot-swap, or a servlet container’s redeploy does, and that part is correct, ordinary behaviour.
static final List<Runnable> LIVE_PLUGINS = new ArrayList<>();
while (true) {
PluginClassLoader loader = new PluginClassLoader(
ClassloaderLeakApplication.class.getClassLoader(), className, classBytes);
Class<?> pluginClass = loader.loadClass(className);
Runnable plugin = (Runnable) pluginClass.getDeclaredConstructor().newInstance();
plugin.run();
// The bug: this reload's plugin instance, Class, and ClassLoader are never
// forgotten.
LIVE_PLUGINS.add(plugin);
}
The bug is LIVE_PLUGINS: a registry that an unload path forgot to clear, which is usually how this shows up in a real codebase — a static “active plugins” list, a metrics map keyed by plugin, a debug registry someone added for introspection and never bounded. Keeping one reload’s plugin instance alive keeps its getClass() alive (a Class object holds a reference to the ClassLoader that defined it), which keeps that specificPluginClassLoader alive, which in turn keeps every other class that loader ever defined alive, whether anything still uses them or not. Full files: ClassloaderLeakApplication.java and PluginClassLoader.java.
Under a 160 MB heap the crash is just as direct as the first three — from 01-oom-console.txt:
classloader-leak: driving repeated plugin reloads to OOM, class=com.ankurm.oomplugin.PluginTask
classloader-leak: reloads=150 livePlugins=150
classloader-leak: reloads=300 livePlugins=300
classloader-leak: reloads=450 livePlugins=450
java.lang.OutOfMemoryError: Java heap space
Dumping heap to docs/output/heap.hprof ...
Heap dump file created [138919083 bytes in 0.140 secs]
Problem Suspect 1 The class com.ankurm.oomdemo.ClassloaderLeakApplication , loaded by
org.springframework.boot.loader.launch.LaunchedClassLoader @ 0xf606bba8 , occupies 123,132,768
(95.35%) bytes. The top consumers of its minimum retained heap are byte[] (1,412 instances
totaling 122,739,984), java.util.concurrent.ConcurrentHashMap (1,404 instances totaling 89,856),
and java.util.concurrent.ConcurrentHashMap$Node[] (936 instances totaling 74,880).
Why the plugin class is a committed .class resource, not an ordinary source file.PluginTask.java lives in plugin-src/, outside the module’s own Maven compilation unit, and is compiled once by build-plugin.sh into a resource the app loads as raw bytes. That is deliberate: the whole demo depends on handing the same class bytes to a differentClassLoader on every reload, which only makes sense if those bytes exist independently of the application’s own classpath, where the JVM’s built-in loader would have already defined and cached the class once and for all.
In this demo each reload’s heap cost is PluginTask’s own 256 KB field, so the crash is a heap exhaustion, not a Metaspace one — worth being precise about, since the two get conflated. In application servers and plugin hosts the same retention graph typically shows up as both: live instances exhaust the heap while the classes and bytecode metadata those classloaders pinned exhaust Metaspace at the same time, which is why repeated java.lang.OutOfMemoryError: Metaspace after every redeploy is the other common face of this exact bug.
ClassLoader javadoc, for the parent-delegation and class-identity rules this pattern depends on.
Pattern five: the queue with no bottom
unbounded-queue-leak is the most common of the five in real systems, because it is the one that arrives disguised as perfectly ordinary producer/consumer code — right up until the consumer falls behind.
// The bug: no capacity bound. new LinkedBlockingQueue<>(10_000) would turn this
// whole failure mode into a put() that blocks instead of an OutOfMemoryError.
LinkedBlockingQueue<byte[]> queue = new LinkedBlockingQueue<>();
Thread consumer = new Thread(() -> {
while (true) {
queue.take();
Thread.sleep(50); // the slow downstream call
}
}, "slow-consumer");
consumer.start();
long payloadSize = 512 * 1024; // 512 KB per message
while (true) {
queue.put(new byte[(int) payloadSize]);
}
new LinkedBlockingQueue<>() with no capacity argument is unbounded. It will accept elements until the heap is gone, full stop — there is no backpressure, no warning, nothing. In a real system this is a fast producer thread and a slow consumer thread (a downstream HTTP call, a database write, anything with latency), and the queue between them quietly becomes the de facto allocator for the entire payload stream the moment the consumer’s rate drops below the producer’s, even for a few minutes during a deploy or a downstream blip. Full file: UnboundedQueueLeakApplication.java.
Heap dies fast with 512 KB messages, and this time the crash arrives inside the producer’s own call to put(), with a clean stack trace — from 01-oom-console.txt:
unbounded-queue-leak: driving an unbounded queue to OOM, consumer sleeps 50ms/item
java.lang.OutOfMemoryError: Java heap space
Dumping heap to docs/output/heap.hprof ...
Heap dump file created [175828534 bytes in 0.250 secs]
Caused by: java.lang.OutOfMemoryError: Java heap space
at java.base/java.util.concurrent.LinkedBlockingQueue.put(LinkedBlockingQueue.java:329)
at com.ankurm.oomdemo.UnboundedQueueLeakApplication.lambda$driveLeak$0(UnboundedQueueLeakApplication.java:53)
at com.ankurm.oomdemo.UnboundedQueueLeakApplication$$Lambda/0x000000000b247688.run(Unknown Source)
and MAT’s suspect, from 02-mat-leak-suspects.txt, names the queue directly — the only one of the five patterns where MAT’s top suspect is the actual collection type rather than the application class wrapping it:
Problem Suspect 1 One instance of java.util.concurrent.LinkedBlockingQueue loaded by <system
class loader> occupies 79,698,088 (92.41%) bytes. The top consumers of its minimum retained heap
are byte[] (152 instances totaling 79,694,208), java.util.concurrent.LinkedBlockingQueue$Node
(153 instances totaling 3,672), and java.util.concurrent.locks.ReentrantLock$NonfairSync (2
instances totaling 64).
The fix really is the one-line comment in the source: pass a capacity to the constructor. Spring‘s own ThreadPoolTaskExecutor has the identical default — setQueueCapacity defaults to unbounded unless you set it, which is exactly what makes a thread pool’s backlog grow without limit under sustained overload instead of rejecting or blocking. Tuning that default alongside heap and GC flags for a production Spring Boot deployment is covered end to end in JVM Flags for Spring Boot in Containers.
LinkedBlockingQueue javadoc — the single-argument constructor is the fix, two lines down from the one this demo calls.
What the defaults do not tell you
Two defaults make this entire workflow harder than it needs to be, and both bite before you ever get to open MAT.
First: -XX:+HeapDumpOnOutOfMemoryError is not on by default. Every transcript in this article exists because the five run scripts pass it explicitly. Without it, a real production OutOfMemoryError leaves you with nothing but a log line — the exact crash this whole article exists to analyze, with no dump to analyze it from. It costs nothing at idle and should be set everywhere, alongside -XX:HeapDumpPath pointed at a volume with enough room: these five dumps ran 138–300 MB at a 160–220 MB heap cap, and a multi-gigabyte production heap can leave a dump several times that size.
Second, and the one that causes more confusing production incidents than any leak pattern in this article: -Xmx with no value, and a container memory limit, are two completely different numbers, and a JVM that overruns the second one never gets the chance to throw OutOfMemoryError at all.
A JVM with no explicit -Xmx picks a default heap cap from the visible memory (historically a quarter of physical RAM; modern JDKs are cgroup-aware and use a fraction of the container’s limit instead, via -XX:MaxRAMPercentage), but the heap is never the JVM’s only memory consumer — thread stacks, Metaspace, the JIT’s code cache, direct ByteBuffers, and the GC’s own bookkeeping all live outside -Xmx and still count against the container’s hard memory limit. Leave too little headroom between -Xmx and that limit and the kernel’s cgroup controller kills the whole process with SIGKILL the instant total resident memory crosses the line — before the JVM’s own heap ever technically fills up, before HeapDumpOnOutOfMemoryError gets a chance to run, before a single log line gets flushed. Killed and exit code 137 is everything you get.
Set -Xmx explicitly, below the container limit, with headroom. Letting the JVM’s cgroup-aware default pick a heap size for you removes your only lever for keeping total RSS under the limit, and it is the difference between crashing with an analyzable dump and crashing with nothing at all. The full flag set for this — heap, Metaspace, container limits, and what each one defaults to on JDK 25 and 27 — is in JVM Flags for Spring Boot in Containers.
On Kubernetes specifically, the pod’s own memory limit is the cgroup limit the paragraph above describes, and a pod killed this way reports OOMKilled in its status with no JVM-side evidence at all — Deploying Spring Boot 4 on Kubernetes covers sizing that limit and the liveness-probe interaction that often makes an OOMKilled loop look like a different failure entirely.
Five different source bugs, five different objects in the dominator tree, and one mental model underneath all of them — here is where each one actually hides:
Pattern
What MAT’s dominator tree names
One-line fix
Static cache, no eviction
the class holding the static field
bound the cache, or stop making it static
ThreadLocal on a pooled thread
the thread / executor, not ThreadLocal itself
call remove() when the task ends
Listener never unregistered
the singleton list’s backing array
remove the listener on “close”
Classloader never released
the class holding the registry, by way of its instances
stop retaining instances after a reload
Unbounded queue
the queue instance directly
give the constructor a capacity
Should you reach for Eclipse MAT at all?
Honestly, not for a first guess. If the leak matches one of the five shapes above and you can point at the suspect field by reading the diff that shipped the week things got slow, fix it and move on — firing up MAT to confirm what a code review already told you is wasted time. A quick jcmd <pid> GC.heap_info or a live-object histogram (jmap -histo:live, or the equivalent endpoint if the app already exposes JFR) across two points in time will usually tell you which class is growing well before you need a full dump, and an APM’s heap trend graph does the same thing continuously, for free, without ever stopping the JVM. Java Flight Recorder and Mission Control covers that lower-overhead, always-on alternative in detail, and it is where most leak investigations should actually start.
Reach for a heap dump and MAT specifically when one of two things is true: the crash already happened in production and HeapDumpOnOutOfMemoryError left you a dump with no other leads, or you have a vague, slow “memory climbs over days” symptom and no hypothesis at all. That is the situation the dominator tree earns its keep in — it will name the retaining object with a percentage attached, which a histogram alone will not, and every one of the five Problem Suspect paragraphs above did that correctly on the first try, with no manual graph-walking required.
Further reading
Companion repository: asmhatre/spring-boot-demo-oom — all five modules, every transcript this article quotes, and the scripts that regenerate them from scratch.
No Comments yet!