Add spring-batch-partitioning: manager/worker partitioning, gridSize semantics, rejected-partition recovery, and a real 10M-row scaling sweep
This commit is contained in:
@@ -0,0 +1,82 @@
|
||||
# 9. JobExecutionAlreadyRunningException, forever — and Spring Batch 6.0's `recover()`
|
||||
|
||||
[← Previous](08-restart-reruns-only-the-failed-partition.md) | [README](../README.md) | [Next: Scaling sensitivity to data size →](10-scaling-sensitivity-to-data-size.md)
|
||||
|
||||
[Chapter 7](07-the-rejectedexecutionexception.md) left `JobExecution` 1 with three worker
|
||||
`StepExecution`s permanently parked at `STARTING`. This chapter is what trying to move on from
|
||||
that actually looks like.
|
||||
|
||||
## The job cannot be restarted
|
||||
|
||||
The same command that would normally restart a failed job, run again against the same
|
||||
`shardsDir`:
|
||||
|
||||
```console
|
||||
$ java -jar target/spring-batch-partitioning-1.0.0.jar \
|
||||
--partition.shards-dir=./data/shards-small --partition.grid-size=4 --partition.pool-core-size=4 --partition.pool-max-size=4
|
||||
```
|
||||
|
||||
throws, every single time:
|
||||
|
||||
```
|
||||
Caused by: org.springframework.batch.core.launch.JobExecutionAlreadyRunningException: A job execution
|
||||
for this job is already running: JobExecution: id=1, version=3, ... status=FAILED, exitStatus=exitCode=FAILED;...
|
||||
```
|
||||
|
||||
Full transcript: [`docs/output/07-restart-throws-alreadyrunning.txt`](output/07-restart-throws-alreadyrunning.txt).
|
||||
Read the exception's own embedded string closely: it says `status=FAILED` in the same message that
|
||||
claims the execution is "already running." `SimpleJobOperator`'s running-check is not "is the
|
||||
`JobExecution` status `FAILED`" — it is closer to "does this `JobInstance` have any
|
||||
`StepExecution` that has not reached a terminal status," and the three orphaned
|
||||
`STARTING`/`EXECUTING` worker steps from chapter 7 satisfy that condition indefinitely. There is no
|
||||
number of retries that fixes this on its own. The job is not failed. It is stuck.
|
||||
|
||||
## `JobOperator#recover`, new in Spring Batch 6.0
|
||||
|
||||
6.0 added exactly the operation this situation calls for:
|
||||
`JobOperator.recover(JobExecution)`. [`RecoveryRunner`](../src/main/java/com/ankurm/batchpartition/runner/RecoveryRunner.java)
|
||||
(active under the `recover` Spring profile, `@Order(0)` so it runs before the normal launch logic)
|
||||
fetches the stuck execution via `JobExplorer` and calls it:
|
||||
|
||||
```console
|
||||
$ java -jar target/spring-batch-partitioning-1.0.0.jar --spring.profiles.active=recover \
|
||||
--recover.job-execution-id=1 --partition.shards-dir=./data/shards-small \
|
||||
--partition.grid-size=4 --partition.pool-core-size=4 --partition.pool-max-size=4
|
||||
|
||||
RECOVER: before -> status=FAILED
|
||||
RECOVER: step=ordersWorkerStep:partition3 status=STARTING
|
||||
RECOVER: step=ordersWorkerStep:partition2 status=STARTING
|
||||
RECOVER: step=ordersWorkerStep:partition1 status=COMPLETED
|
||||
RECOVER: step=ordersWorkerStep:partition0 status=STARTING
|
||||
RECOVER: after -> status=FAILED
|
||||
RECOVER: step=ordersWorkerStep:partition3 status=FAILED
|
||||
RECOVER: step=ordersWorkerStep:partition2 status=FAILED
|
||||
RECOVER: step=ordersWorkerStep:partition1 status=COMPLETED
|
||||
RECOVER: step=ordersWorkerStep:partition0 status=FAILED
|
||||
JOB FINISHED: id=33 status=COMPLETED exitCode=COMPLETED
|
||||
```
|
||||
|
||||
Full transcript: [`docs/output/08-recover-then-restart.txt`](output/08-recover-then-restart.txt).
|
||||
`recover()` force-closes every non-terminal `StepExecution` under the given `JobExecution` to
|
||||
`FAILED` — nothing else. The already-`COMPLETED` `partition1` is left exactly as it was.
|
||||
Immediately afterward, in the same JVM, `OrderIngestRunner`'s ordinary `jobOperator.start(...)`
|
||||
call runs against the same `shardsDir` and this time succeeds, producing `JobExecution` 33 —
|
||||
which [chapter 8](08-restart-reruns-only-the-failed-partition.md) shows re-executed only the three
|
||||
partitions `recover()` had just marked `FAILED`.
|
||||
|
||||
## The takeaway
|
||||
|
||||
A `RejectedExecutionException` (or any other exception that reaches the executor *before* a
|
||||
`StepExecution` starts, rather than during it) is a different failure class than an ordinary step
|
||||
failure, and it needs a different remedy: `recover()`, then restart. Sizing the pool correctly in
|
||||
the first place ([chapter 3](03-what-gridsize-actually-controls.md)) avoids the situation
|
||||
entirely; `recover()` is what to reach for when a job in production has already gotten into it.
|
||||
|
||||
## Going deeper
|
||||
|
||||
- `JobOperator`'s full interface, including `recover`, `abandon`, and `stop`:
|
||||
[`docs/output/03-package-repackaging-javap.txt`](output/03-package-repackaging-javap.txt) has the
|
||||
decompiled method list this chapter's claims were checked against.
|
||||
- The stuck state this chapter recovers from: [chapter 7](07-the-rejectedexecutionexception.md).
|
||||
|
||||
[Next: Scaling sensitivity to data size →](10-scaling-sensitivity-to-data-size.md)
|
||||
Reference in New Issue
Block a user