Adds a JUnit test suite (GetVsGetReferenceTest, MergeRefreshTest, OptimisticLockTest, IdentityBatchTest, SequenceBatchTest, AllocationSizeSweepTest, BatchSizeSweepTest) so every surprising behavior described in the three companion posts has a reproducible test, alongside the original CommandLineRunner scenarios. Rewrites all three doc chapters and the README around the new experiments: the get()/getReference() same-session matrix, the merge()/refresh() experiments (including exactly when OptimisticLockException surfaces and a corrected LAZY-plus- cascade merge() result), and two new sweeps (allocationSize, batch_size) for batch inserts.
9.2 KiB
02 — merge() vs refresh()
← Previous: 01 — get() vs getReference() | Next: 03 — Inserting objects →
Backs ankurm.com: merge() vs refresh().
Test classes: MergeRefreshTest,
OptimisticLockTest.
Entity: Book — note the real
@Version column and the notes LAZY collection, both load-bearing for the experiments below.
./mvnw -Dtest=MergeRefreshTest,OptimisticLockTest test
Raw captured output: docs/output/merge-vs-refresh-tests.txt.
(The docs/output/merge-vs-refresh.txt file, unchanged, is the older CommandLineRunner
transcript behind the mergerefresh profile mentioned in the README — a different, narrower
scenario than the three experiments below.)
Contract vs observation
The textbook framing is "merge() pushes Java state to the database, refresh() pulls database
state into Java — opposite directions of the same kind of operation." That's accurate as a
description of data flow and useless as a guide to which one is dangerous. The three experiments
below are about the part the contract doesn't specify: what each method does when the state it's
holding is already stale, which is precisely the situation both exist to handle.
Reframing: this is about entity state, not "data consistency"
merge() and refresh() don't care about your application's notion of consistency. They care
about exactly one thing: what identity state (@Version value, or an unflushed field on a managed
instance) the object handed to them holds at the moment they're called. Everything below follows
from that, entity-state mechanics, not from anything data-consistency-flavored.
Experiment 1 — merge() of a detached instance
Config: a Book row exists in the database. A session already holds its own managed
instance of that row (via get()), before merge() is ever called on a separately-detached copy
of the same row.
Expected: merge() returns some object representing the updated state.
Observed: merge() returns the exact, identity-equal managed instance the session already had
— not a new object, and not the detached instance passed in:
Book result = session.merge(detached);
assertThat(result).isSameAs(managed); // true
assertThat(result).isNotSameAs(detached); // true
This is the precise version of "merge() returns a managed copy" — it's not just a managed
copy, it's the one instance this persistence context has already committed to tracking for this
row, reused rather than replaced.
Experiment 2 — optimistic-lock conflict: WHEN does it surface?
Config: a detached Book instance holds version=0. A second, independent session has since
updated the same row and committed, advancing the database to version=1. The stale detached
instance is then edited and merged.
Expected (the common but imprecise claim): "merge() throws OptimisticLockException."
Observed, precisely: it does — and specifically at the merge() call itself, not at
flush() and not at tx.commit():
OptimisticLockException surfaced directly from the merge() call.
exception: jakarta.persistence.OptimisticLockException: Row was already updated or deleted by
another transaction for entity [com.ankurm.hibernatedemo.model.Book with id '4']
The test doesn't assume this — it wraps session.merge(detached) in a try/catch first and only
falls through to asserting the exception at commit() if merge() itself didn't throw, logging
which path actually happened. On 7.4.1.Final, merge() re-selects the row as part of copying
state and compares versions right there, before any flush is even scheduled — so the failure is as
early as it can possibly be. This matters in code that wraps merge() calls expecting the
exception only at commit time: on this version, it never gets that far.
Experiment 3 — merge() with a LAZY collection under CascadeType.MERGE
Config: Book.notes is FetchType.LAZY and cascades MERGE. A Book is loaded and detached
without ever touching .getNotes() — the collection proxy is confirmed uninitialized before
detachment. The detached instance is edited and merged.
Expected (the naive, plausible-sounding claim): "an unfetched LAZY collection is never
touched by merge(), since it was never loaded in the first place."
Observed: the opposite. merge() initializes the collection anyway:
Book merged = session.merge(detached);
assertThat(Hibernate.isInitialized(merged.getNotes())).isTrue(); // true -- NOT false
The reason is cascading itself: CascadeType.MERGE on notes means merging the parent requires
merging each element of that collection too, and Hibernate cannot cascade to elements it hasn't
loaded — so it loads them first. Remove cascade = CascadeType.MERGE from Book.notes and
re-run this exact test and the result flips: the collection stays uninitialized, because nothing
requires Hibernate to look at it. Cascading, not laziness, decides whether merge() touches an
unfetched collection.
Experiment 4 — refresh() silently discards an unflushed edit
Config: the USER_EDIT / ADMIN_EDIT scenario. A Book.status row starts at USER_EDIT. A
separate admin process loads the row, sets status=ADMIN_EDIT, and commits. A second session then
loads the row (now ADMIN_EDIT in the database), makes a local, unflushed edit back to
USER_EDIT, and calls session.refresh() on the managed instance.
Observed:
managed.setStatus("USER_EDIT"); // local edit, never sent to the database
session.refresh(managed);
assertThat(managed.getStatus()).isEqualTo("ADMIN_EDIT"); // the local edit is gone
No exception. No warning. refresh() re-runs the SELECT and overwrites every field on the
managed instance with what the database currently holds — including the field holding the pending
edit that was never flushed. The edit isn't rejected; it's erased.
What surprised me building this
Going in, the plan was to demonstrate "merge() can silently overwrite concurrent changes" as
the headline risk — that's the framing most write-ups use, and it's the one the article originally
carried. Running Experiments 2 and 4 back to back showed the opposite. With a real @Version
column in place, merge() is the one that refuses to write a stale change — loudly, at the
earliest possible point. refresh() is the one that destroys data without a sound, and it does it
to an edit that was never even sent to the database. The risk isn't "which method can overwrite
the database" — both can, that's their job. It's "which one fails loudly when the state it's
holding is stale," and on a versioned entity that's refresh(), not merge() — backwards from
how the pairing is usually described. Strip the @Version column out and Experiment 2 flips: an
unversioned merge() would apply the stale write without complaint. The column isn't incidental
to the result; it's the entire reason the result comes out this way.
The LAZY-collection result (Experiment 3) was the other correction: the plan going in was to
show that unfetched LAZY state is never touched by merge(). It is, specifically because of the
cascade — a fact only visible by running the on/off comparison rather than asserting the more
intuitive-sounding half of it.
Decision tree
Do you have a DETACHED instance you want written to the database?
├─ Yes → merge()
│ Does the entity carry @Version?
│ ├─ Yes → a stale write throws OptimisticLockException AT THE merge() CALL — loud, safe
│ └─ No → a stale write silently overwrites the current row — no different from update()
│
└─ No, you have a MANAGED instance and want it to reflect the current database row
→ refresh()
Does it have an unflushed local edit?
├─ Yes → that edit is silently discarded, no exception — refresh() is NOT reversible
└─ No → refresh() is a safe, ordinary re-read
Pessimistic locking, briefly
Pessimistic locking (LockModeType.PESSIMISTIC_WRITE, issuing SELECT ... FOR UPDATE) is the
other tool for the same underlying problem — it prevents the conflict from ever existing rather
than detecting it after the fact. Reach for it only when the retry cost of an
OptimisticLockException is genuinely unacceptable (real-time seat/ticket reservation, high
per-row contention); it holds a database lock for the duration of the transaction, which is the
wrong tradeoff for the common case of a REST API with real user think-time between load and save.
@Version optimistic locking is the sane default; this repo doesn't carry a dedicated scenario for
pessimistic locking because there's no surprising runtime behavior to verify here beyond "the lock
is held until commit," which the database's own documentation already states correctly.
← Previous: 01 — get() vs getReference() | Next: 03 — Inserting objects →