Releases: ceruleane/gpusched
Releases · ceruleane/gpusched
Release list
v0.4.0
Changelog
0.4.0
- Cancellation: add the bare
canceltoken to a job's attribute block to
stop it if running (SIGTERM -> SIGKILL after grace, plus a straggler sweep
so TERM-trapping descendants don't outlive the job) or to prevent it from
starting if pending. Identity hashes the command text, so the token targets
the same job. Reported separately from failures; never consumes OOM
retries; does not affect the scheduler exit code. Deleting a running line
remains a deliberate no-op. - Orphan double-run guard: launches are journaled with their pgid; on
restart, a still-alive previous attempt blocks re-dispatch with a clear
warning, and a dead one is marked interrupted and re-queued. Closes the
silent-corruption case of two copies of one job writing the same outputs.
0.3.0
- Live-editable queue: the jobs file is user-owned, re-read every poll;
append / delete / reorder pending lines mid-run. Pending = in-file and
neither running nor terminal in the journal. - Journal (
journal.jsonl): per-job identities, attempts, outcomes.
Resume-after-restart falls out of it;--freshresets. - CUDA-OOM-aware retry (
[retries=N]/--oom-retries) with the
declaration auto-bumped to ~1.25x of observed peak; non-OOM failures
never consume retries. - Opt-in per-job walltime (
[timeout=...]): SIGTERM, then SIGKILL after a
grace period. No heuristic hang detection by design. - Per-job average device utilization in completion reports.
- Live status board rendered to
<log_dir>/status.txtevery poll. --watch: keep running after drain, picking up appended lines.- Parser hardening: an unterminated leading
[...]block is a parse error
(previously it fell through and was executed as a shell command); a
malformed mid-edit jobs file keeps the last good queue.
0.2.0
- High-water-mark scheduling: external processes are held to their observed
per-GPU peaks (buffered by--spike-buffer) until they exit, so momentary
VRAM troughs are not treated as packable space; a scheduled job that
exceeds its declaration has its budget escalated to its buffered peak. - Idle detection for undeclared jobs judged on effective (peak-aware)
external usage.
0.1.0
- Initial release: VRAM-declaration-based placement with launch-race
reservations, packing,--exclusive, multi-GPU jobs (gpus=N),
per-job VRAM attribution by process group, immediate under-declaration
warnings, completion-time over-declaration reports, backfill, fail-fast
infeasibility, simulated backends and--simdry-run mode.
v0.3.0
Changelog
0.3.0
- Live-editable queue: the jobs file is user-owned, re-read every poll;
append / delete / reorder pending lines mid-run. Pending = in-file and
neither running nor terminal in the journal. - Journal (
journal.jsonl): per-job identities, attempts, outcomes.
Resume-after-restart falls out of it;--freshresets. - CUDA-OOM-aware retry (
[retries=N]/--oom-retries) with the
declaration auto-bumped to ~1.25x of observed peak; non-OOM failures
never consume retries. - Opt-in per-job walltime (
[timeout=...]): SIGTERM, then SIGKILL after a
grace period. No heuristic hang detection by design. - Per-job average device utilization in completion reports.
- Live status board rendered to
<log_dir>/status.txtevery poll. --watch: keep running after drain, picking up appended lines.- Parser hardening: an unterminated leading
[...]block is a parse error
(previously it fell through and was executed as a shell command); a
malformed mid-edit jobs file keeps the last good queue.
0.2.0
- High-water-mark scheduling: external processes are held to their observed
per-GPU peaks (buffered by--spike-buffer) until they exit, so momentary
VRAM troughs are not treated as packable space; a scheduled job that
exceeds its declaration has its budget escalated to its buffered peak. - Idle detection for undeclared jobs judged on effective (peak-aware)
external usage.
0.1.0
- Initial release: VRAM-declaration-based placement with launch-race
reservations, packing,--exclusive, multi-GPU jobs (gpus=N),
per-job VRAM attribution by process group, immediate under-declaration
warnings, completion-time over-declaration reports, backfill, fail-fast
infeasibility, simulated backends and--simdry-run mode.