1
0

Add hibernate-demo: get() vs load(), merge() vs refresh(), inserting objects (Hibernate 7.4.1.Final + Spring Boot 4.1.0)

Adds a JUnit test suite (GetVsGetReferenceTest, MergeRefreshTest, OptimisticLockTest,
IdentityBatchTest, SequenceBatchTest, AllocationSizeSweepTest, BatchSizeSweepTest) so every
surprising behavior described in the three companion posts has a reproducible test, alongside
the original CommandLineRunner scenarios. Rewrites all three doc chapters and the README around
the new experiments: the get()/getReference() same-session matrix, the merge()/refresh()
experiments (including exactly when OptimisticLockException surfaces and a corrected LAZY-plus-
cascade merge() result), and two new sweeps (allocationSize, batch_size) for batch inserts.
This commit is contained in:
2026-08-26 17:26:04 +00:00
commit 7b06d653e4
43 changed files with 3053 additions and 0 deletions

View File

@@ -0,0 +1,198 @@
# 03 — Hibernate 7 batch inserts: proving batching is working
[← Previous: 02 — merge() vs refresh()](02-merge-vs-refresh.md) | [Back to README →](../README.md)
Backs [ankurm.com: Hibernate 7 batch inserts](https://ankurm.com/mastering-hibernate-7-the-ultimate-guide-to-inserting-objects-efficiently/).
Test classes: [`IdentityBatchTest`](../src/test/java/com/ankurm/hibernatedemo/IdentityBatchTest.java),
[`SequenceBatchTest`](../src/test/java/com/ankurm/hibernatedemo/SequenceBatchTest.java),
[`AllocationSizeSweepTest`](../src/test/java/com/ankurm/hibernatedemo/AllocationSizeSweepTest.java),
[`BatchSizeSweepTest`](../src/test/java/com/ankurm/hibernatedemo/BatchSizeSweepTest.java).
```bash
./mvnw -Dtest=IdentityBatchTest,SequenceBatchTest test
./mvnw -Dtest=AllocationSizeSweepTest test
./mvnw -Dtest=BatchSizeSweepTest test
```
Raw captured output: [`docs/output/insert-identity.txt`](output/insert-identity.txt),
[`docs/output/insert-sequence.txt`](output/insert-sequence.txt),
[`docs/output/allocation-and-batch-size-sweeps.txt`](output/allocation-and-batch-size-sweeps.txt).
## The discovery that opens this chapter: IDENTITY silently disables batching
Two entities, [`WidgetIdentity`](../src/main/java/com/ankurm/hibernatedemo/model/WidgetIdentity.java)
and [`WidgetSequence`](../src/main/java/com/ankurm/hibernatedemo/model/WidgetSequence.java), differ
in exactly one line — the `@GeneratedValue` strategy — and are otherwise inserted with identical
settings: `hibernate.jdbc.batch_size=25`, `hibernate.order_inserts=true`. 30 rows each, same
transaction shape, `hibernate.generate_statistics=true` reading the real counts.
## Results, up front
| | `entityInsertCount` | `prepareStatementCount` | Batching actually happening? |
|---|---|---|---|
| `WidgetIdentity` (`GenerationType.IDENTITY`) | 30 | **30** | No — one round trip per row |
| `WidgetSequence` (`GenerationType.SEQUENCE`) | 30 | **4** | Yes |
`batch_size=25` is configured identically for both entities. It does nothing at all for
`IDENTITY``prepareStatementCount` equals `entityInsertCount` exactly, meaning every insert is
its own round trip. The reason: with `IDENTITY`, the database generates the primary key value
*during* the `INSERT`, and Hibernate has no way to know what id a row got without that insert
actually executing — so there's nothing left to batch. `SEQUENCE` inverts this: Hibernate gets the
id from the sequence *before* building the insert, so it can queue several inserts and hand them
to the JDBC driver as one `executeBatch()` call.
## Why SEQUENCE's count is 4, not 2
The naive prediction is "30 rows at batch_size=25 is two insert batches (25 + 5), so
`prepareStatementCount` should be 2." Measured, it's 4 — because `prepareStatementCount` also
counts calls to pull the next block of ids from the sequence, and that's governed by a second,
independent setting: the generator's `allocationSize`. `WidgetSequence` sets `allocationSize=25`
to match `batch_size`, so the first 25 ids come from one sequence call and the remaining 5 force a
second — 2 insert batches + 2 sequence calls = 4. `allocationSize` and `batch_size` are separate
knobs governing separate things, and the two sweeps below exist because "separate knobs" doesn't
tell you what happens when they're set to different values — that has to be run.
## Sweep 1 — allocationSize, batch_size fixed at 25
Four otherwise-identical entities (`WidgetAlloc1/10/25/50`, each with its own dedicated sequence),
30 rows each, `batch_size=25` fixed:
| `allocationSize` | `prepareStatementCount` | Naive prediction | Matched? |
|---|---|---|---|
| 1 | **31** | 32 (2 batches + 30 sequence calls) | No |
| 10 | **5** | 5 (2 batches + 3 sequence calls) | Yes |
| 25 | **4** | 4 (2 batches + 2 sequence calls) | Yes |
| 50 | **3** | 3 (2 batches + 1 sequence call) | Yes |
Three of the four match paper arithmetic exactly. `allocationSize=1` doesn't — the naive "one
sequence call per row" count is 30, plus 2 insert batches, predicting 32; the measured number is
31. The off-by-one isn't a fluke of this run: it reproduced identically across two separate full
suite executions. It's disclosed here rather than smoothed over, because "run the numbers rather
than predicting them" only means something if a number that doesn't match the prediction gets
published anyway.
## Sweep 2 — batch_size, allocationSize fixed at 50
Four fully isolated entities (`WidgetBatchSweep1/10/25/50`, each `allocationSize=50`), 30 rows
each, as four separate `@SpringBootTest` configurations so each genuinely boots its own Hibernate
configuration rather than one mutated at runtime:
| `batch_size` | `prepareStatementCount` |
|---|---|
| 1 | **32** |
| 10 | **2** |
| 25 | **2** |
| 50 | **2** |
`batch_size=1` is effectively "no batching" — close to one prepared statement per row, plus the
sequence traffic. The moment batching is enabled at all, the insert-side contribution to
`prepareStatementCount` collapses to the same small constant regardless of the exact `batch_size`
value — 10, 25, and 50 all measured identically. That does **not** mean `batch_size` stops
mattering: it still governs how many rows go into each `executeBatch()` call at the JDBC driver
level, which is real and documented — it just isn't a distinction this particular Hibernate
statistic can see once batching is switched on at all. Reading `prepareStatementCount` answers
"is batching happening," not "how big are the batches."
**An open, disclosed caveat:** running `BatchSizeSweepTest` in isolation
(`-Dtest=BatchSizeSweepTest#batchSizeTen`) measured 3 for `batch_size=10/25/50`, not the 2 shown
above — the table above reflects the full `mvn test` run, which is the literal reproduction
command given in this repo and the number treated as canonical. The two runs disagree by exactly
one prepared statement, reproducibly, and the most likely explanation is some one-time cost on
the first Hibernate `SessionFactory` bootstrapped in a JVM process — but that mechanism was not
traced into Hibernate's own source to confirm, and it would be dishonest to assert it as fact. Run
both ways yourself; they're one flag apart.
## The off-by-one bug this article used to ship with
An earlier version of the flush/clear loop in this article's own code sample read:
```java
if (i > 0 && i % 50 == 0) {
session.flush();
session.clear();
}
```
`i` is the 0-based loop index (`for (int i = 0; i < users.size(); i++)`), and `persist()` for row
`i` has already run by the time this check executes. `i > 0 && i % 50 == 0` is true at
`i = 50, 100, 150, ...` — but by the time `i` reaches 50, rows at index `0` through `50` have
already been persisted, which is **51 rows**, not 50. Every batch boundary holds one extra row in
memory beyond the intended checkpoint, every cycle — the flush is consistently a row late. The
fix:
```java
if ((i + 1) % 50 == 0) {
session.flush();
session.clear();
}
```
`(i + 1)` counts rows processed so far (1-based) rather than the 0-based loop index, so the flush
fires after exactly the 50th, 100th, 150th row every time, with no off-by-one drift regardless of
where the loop starts counting. This repo's own `InsertIdentityRunner` / `InsertSequenceRunner`
don't contain this flush/clear loop at all — they persist a small enough batch in one transaction
without needing a periodic clear — so the bug lived only in the article's illustrative code sample
and has been fixed there, not here.
## Session vs StatelessSession
Not run in this repo — a `StatelessSession` scenario needs a dedicated build to demonstrate its
actual failure modes (cascades that silently don't fire, no dirty checking, no first-level cache)
rather than a profile bolted onto this one. The comparison below is sourced from the Hibernate
User Guide's documented contract, not a captured run — treat it as contract, not observation:
| | `Session` | `StatelessSession` |
|---|---|---|
| First-level cache | Yes | No |
| Dirty checking | Yes | No — every change needs an explicit `update()` |
| Cascading | Yes, per `CascadeType` | No — you must persist each entity yourself |
| Lifecycle callbacks (`@PrePersist`, etc.) | Yes | No |
| Batching behavior | Governed by `hibernate.jdbc.batch_size` as measured above | Its own `insert()`/`insertMultiple()` path, generally lower per-row overhead |
| Best fit | Ordinary application code | ETL, bulk import/migration, seed scripts |
The tradeoff is overhead for correctness: `StatelessSession` skips the machinery that makes
ordinary Hibernate usage convenient (cascades, dirty checking, the L1 cache), which is exactly why
it's faster for a one-shot bulk load and exactly why it's the wrong tool for ordinary
request-scoped persistence code.
## Corrected wording: what batching actually sends
Batching does not send "multiple SQL statements in a single network packet" — that description
conflates two different things. What actually happens: JDBC's `PreparedStatement.addBatch()` /
`executeBatch()` groups several parameter sets for the *same* prepared statement and sends them in
one client-to-server exchange, avoiding a full round-trip per row. Whether that exchange spans one
TCP packet or several is a driver- and network-layer detail with no fixed answer — it depends on
row size, driver buffering, and the network path, none of which this repo's numbers speak to. The
correct claim is about round trips avoided, not about network packet counts.
## Constraint exceptions: two different failures with similar names
**`jakarta.validation.ConstraintViolationException`** (Bean Validation, `jakarta.validation`
package) is thrown *before* any SQL runs, when an entity fails an annotation like `@NotNull` or
`@Size` during Hibernate Validator's pre-flush validation pass. No database round trip happens at
all in this case.
**`org.hibernate.exception.ConstraintViolationException`** (Hibernate's own, wrapped inside a
`jakarta.persistence.PersistenceException`) is thrown *after* SQL runs and the database itself
rejects the statement — a unique index, a foreign key, or a check constraint failing at the
database level.
Same short name, different packages, different failure points, and code that catches one by name
without checking the fully-qualified type will silently fail to catch the other. Catching
`jakarta.persistence.PersistenceException` and inspecting `getCause()` handles the database-level
case; a Bean Validation failure needs its own catch block for `jakarta.validation.ConstraintViolationException`
before that.
## What surprised me building this
The identity-vs-sequence result itself wasn't the surprise — that IDENTITY disables batching is
documented, if you know to look. The surprise was in the sweeps: `allocationSize=1` landing at 31
instead of the paper-arithmetic 32, and the batch_size sweep collapsing to the same constant the
moment batching is enabled at all regardless of the exact value, in a way that changes depending
on whether the test runs alone or as part of the full suite. None of those three things would have
made it into this article from reasoning about the configuration in the abstract — they only show
up by actually running the sweep and being willing to publish a number that didn't match the
prediction.
[← Previous: 02 — merge() vs refresh()](02-merge-vs-refresh.md) | [Back to README →](../README.md)