Restart Mid-Build

Surviving a Controller Restart Mid-Build

The durable-demo job holds a Groovy variable across a 30-second shell loop on agent-1:

durable-demo: Groovy state on both sides of a long sh stepPython
def ticks = 0
node('linux') {
  echo "Groovy state before: ticks=${ticks}"
  sh '''set +x
    for i in $(seq 1 30); do echo "tick $i at $(date +%T)"; sleep 1; done'''
  ticks = 30
}
echo "Groovy state after: ticks=${ticks}"

Six seconds into the loop, the controller container was restarted, a graceful stop (SIGTERM) and a cold start; a minute later the console read:

Restarting the controller in the middle of the buildShell
docker restart l2-jenkins
sleep 60
curl -s -u "admin:$JENKINS_TOKEN" http://localhost:32080/job/durable-demo/3/consoleText
Output
Groovy state before: ticks=0
[Pipeline] sh
+ set +x
tick 1 at 13:48:31
...
tick 7 at 13:48:38
Pausing (shutting down)
Resuming build at Fri Sep 25 13:49:01 UTC 2026 after Jenkins restart
Waiting for reconnection of agent-1 before proceeding with build
Ready to run at Fri Sep 25 13:49:09 UTC 2026
tick 8 at 13:48:39
...
tick 30 at 13:49:01
...
Groovy state after: ticks=30
Finished: SUCCESS

Read the timestamps. The controller was down from 13:48:38 to 13:49:01, yet ticks 8 to 30 were all printed in that window: the loop never stopped, because a sh step is a durable task (Trigger to Executor). On the agent it ran as a detached sh -xe .../durable-demo@tmp/durable-3f303c63/script.sh.copy writing to a log file in that directory. At shutdown the controller saved program.dat and printed "Pausing"; after the restart it deserialized the program, waited 8 seconds for agent-1 to reconnect, re-attached to the durable task, copied the lines it had missed, and resumed the Groovy code with ticks intact.

How much Pipeline writes to disk to make this possible is configurable, globally or per job with properties([durabilityHint('PERFORMANCE_OPTIMIZED')]):

Pipeline durability settings
Durability hint Disk writes After an unclean crash (SIGKILL, power loss)
MAX_SURVIVABILITY (default) Every step, atomically Resumes
SURVIVABLE_NONATOMIC Every step, not atomic Usually resumes
PERFORMANCE_OPTIMIZED Only at graceful shutdown May not resume; the log survives

build.xml records the hint each build ran with (<durabilityHint>MAX_SURVIVABILITY</durabilityHint> here). PERFORMANCE_OPTIMIZED greatly reduces disk I/O and suits short, repeatable builds; keep the default for deployments and anything waiting on input.