# 9. JobExecutionAlreadyRunningException, forever — and Spring Batch 6.0's `recover()` [← Previous](08-restart-reruns-only-the-failed-partition.md) | [README](../README.md) | [Next: Scaling sensitivity to data size →](10-scaling-sensitivity-to-data-size.md) [Chapter 7](07-the-rejectedexecutionexception.md) left `JobExecution` 1 with three worker `StepExecution`s permanently parked at `STARTING`. This chapter is what trying to move on from that actually looks like. ## The job cannot be restarted The same command that would normally restart a failed job, run again against the same `shardsDir`: ```console $ java -jar target/spring-batch-partitioning-1.0.0.jar \ --partition.shards-dir=./data/shards-small --partition.grid-size=4 --partition.pool-core-size=4 --partition.pool-max-size=4 ``` throws, every single time: ``` Caused by: org.springframework.batch.core.launch.JobExecutionAlreadyRunningException: A job execution for this job is already running: JobExecution: id=1, version=3, ... status=FAILED, exitStatus=exitCode=FAILED;... ``` Full transcript: [`docs/output/07-restart-throws-alreadyrunning.txt`](output/07-restart-throws-alreadyrunning.txt). Read the exception's own embedded string closely: it says `status=FAILED` in the same message that claims the execution is "already running." `SimpleJobOperator`'s running-check is not "is the `JobExecution` status `FAILED`" — it is closer to "does this `JobInstance` have any `StepExecution` that has not reached a terminal status," and the three orphaned `STARTING`/`EXECUTING` worker steps from chapter 7 satisfy that condition indefinitely. There is no number of retries that fixes this on its own. The job is not failed. It is stuck. ## `JobOperator#recover`, new in Spring Batch 6.0 6.0 added exactly the operation this situation calls for: `JobOperator.recover(JobExecution)`. [`RecoveryRunner`](../src/main/java/com/ankurm/batchpartition/runner/RecoveryRunner.java) (active under the `recover` Spring profile, `@Order(0)` so it runs before the normal launch logic) fetches the stuck execution via `JobExplorer` and calls it: ```console $ java -jar target/spring-batch-partitioning-1.0.0.jar --spring.profiles.active=recover \ --recover.job-execution-id=1 --partition.shards-dir=./data/shards-small \ --partition.grid-size=4 --partition.pool-core-size=4 --partition.pool-max-size=4 RECOVER: before -> status=FAILED RECOVER: step=ordersWorkerStep:partition3 status=STARTING RECOVER: step=ordersWorkerStep:partition2 status=STARTING RECOVER: step=ordersWorkerStep:partition1 status=COMPLETED RECOVER: step=ordersWorkerStep:partition0 status=STARTING RECOVER: after -> status=FAILED RECOVER: step=ordersWorkerStep:partition3 status=FAILED RECOVER: step=ordersWorkerStep:partition2 status=FAILED RECOVER: step=ordersWorkerStep:partition1 status=COMPLETED RECOVER: step=ordersWorkerStep:partition0 status=FAILED JOB FINISHED: id=33 status=COMPLETED exitCode=COMPLETED ``` Full transcript: [`docs/output/08-recover-then-restart.txt`](output/08-recover-then-restart.txt). `recover()` force-closes every non-terminal `StepExecution` under the given `JobExecution` to `FAILED` — nothing else. The already-`COMPLETED` `partition1` is left exactly as it was. Immediately afterward, in the same JVM, `OrderIngestRunner`'s ordinary `jobOperator.start(...)` call runs against the same `shardsDir` and this time succeeds, producing `JobExecution` 33 — which [chapter 8](08-restart-reruns-only-the-failed-partition.md) shows re-executed only the three partitions `recover()` had just marked `FAILED`. ## The takeaway A `RejectedExecutionException` (or any other exception that reaches the executor *before* a `StepExecution` starts, rather than during it) is a different failure class than an ordinary step failure, and it needs a different remedy: `recover()`, then restart. Sizing the pool correctly in the first place ([chapter 3](03-what-gridsize-actually-controls.md)) avoids the situation entirely; `recover()` is what to reach for when a job in production has already gotten into it. ## Going deeper - `JobOperator`'s full interface, including `recover`, `abandon`, and `stop`: [`docs/output/03-package-repackaging-javap.txt`](output/03-package-repackaging-javap.txt) has the decompiled method list this chapter's claims were checked against. - The stuck state this chapter recovers from: [chapter 7](07-the-rejectedexecutionexception.md). [Next: Scaling sensitivity to data size →](10-scaling-sensitivity-to-data-size.md)