Two fixes that affected any real job:
- The node's heartbeat stopped for the entire duration of a running job, because job execution blocked the poll loop. The scheduler drops a node after 90 seconds without one and the stuck-job reaper fails its running jobs after five minutes, so any job longer than that had its own node declared offline and the work reaped out from under it. Liveness now continues throughout a job.
- Workload images are pulled before the job's runtime budget starts. These images are several GB; previously a first job could spend its entire paid allowance downloading and then be killed at the cap having produced nothing.