You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Cancellation: add the bare cancel token to a job's attribute block to
stop it if running (SIGTERM -> SIGKILL after grace, plus a straggler sweep
so TERM-trapping descendants don't outlive the job) or to prevent it from
starting if pending. Identity hashes the command text, so the token targets
the same job. Reported separately from failures; never consumes OOM
retries; does not affect the scheduler exit code. Deleting a running line
remains a deliberate no-op.
Orphan double-run guard: launches are journaled with their pgid; on
restart, a still-alive previous attempt blocks re-dispatch with a clear
warning, and a dead one is marked interrupted and re-queued. Closes the
silent-corruption case of two copies of one job writing the same outputs.
0.3.0
Live-editable queue: the jobs file is user-owned, re-read every poll;
append / delete / reorder pending lines mid-run. Pending = in-file and
neither running nor terminal in the journal.
Journal (journal.jsonl): per-job identities, attempts, outcomes.
Resume-after-restart falls out of it; --fresh resets.
CUDA-OOM-aware retry ([retries=N] / --oom-retries) with the
declaration auto-bumped to ~1.25x of observed peak; non-OOM failures
never consume retries.
Opt-in per-job walltime ([timeout=...]): SIGTERM, then SIGKILL after a
grace period. No heuristic hang detection by design.
Per-job average device utilization in completion reports.
Live status board rendered to <log_dir>/status.txt every poll.
--watch: keep running after drain, picking up appended lines.
Parser hardening: an unterminated leading [...] block is a parse error
(previously it fell through and was executed as a shell command); a
malformed mid-edit jobs file keeps the last good queue.
0.2.0
High-water-mark scheduling: external processes are held to their observed
per-GPU peaks (buffered by --spike-buffer) until they exit, so momentary
VRAM troughs are not treated as packable space; a scheduled job that
exceeds its declaration has its budget escalated to its buffered peak.
Idle detection for undeclared jobs judged on effective (peak-aware)
external usage.
0.1.0
Initial release: VRAM-declaration-based placement with launch-race
reservations, packing, --exclusive, multi-GPU jobs (gpus=N),
per-job VRAM attribution by process group, immediate under-declaration
warnings, completion-time over-declaration reports, backfill, fail-fast
infeasibility, simulated backends and --sim dry-run mode.