Repository navigation
Releases: llnl/HPC-launcher
Release list
v1.0.6
HPC-launcher v1.0.6
Nested job steps inside existing allocations (#71)
Blocking launches invoked from inside a scheduler allocation (salloc, sbatch,
flux alloc) now run as a job step in the current allocation instead of
accidentally requesting a new one.
- Inherit allocation size: with no job-size flag (
-N/-g/--gpumem-at-least),
the node count is inherited from the enclosing allocation via a new
scheduler-agnostic probe covering Flux, SLURM, and LSF. Explicit flags still
take precedence. - New
-c/--cpus-per-taskflag to give a job or step an explicit CPU
footprint (SLURM--cpus-per-task, Flux--cores-per-task/--cores-per-slot). - SLURM job steps: allocation-selection flags (
--partition,--account,
--reservation) are dropped with a warning inside an allocation; node requests
exceeding the allocation are rejected upfront; and--overlapis added by
default so concurrent steps don't hang on busy nodes (omitted with
--cpus-per-taskor--exclusive, removable with-x ~--overlap). - Non-blocking (
sbatch/flux batch) submissions are unchanged and still create
a new job.
Includes 17 new scheduler-free tests (tests/nested_job_step_test.py) and
updated CLI documentation.
v1.0.5
Release v1.0.5
A large bug-fix and hardening release driven by two rounds of deep code
review, with end-to-end validation on El Capitan-class (MI300A/Flux)
systems.
GPU assignment and the torchrun trampoline
- Workers no longer initialize HIP/CUDA at import time with all GPUs
visible, which pinned every rank on a node to physical GPU 0; each
rank now sees only its granted devices, and the memory-fraction cap
applies to the device actually selected. - The trampoline puts the user's own code on sys.path (sibling imports
work under --local, subdirectories, and custom launch dirs), destroys
the process group on every exit path, publishes rank identity itself,
and only passes device_id to init_process_group for accelerators, so
CPU/gloo jobs start again.
Correctness of generated launches
- Job exit codes propagate through Scheduler.launch to the CLI exit
status; SIGINT kills the scheduler child instead of orphaning it. - User-controlled values are shell-quoted in generated scripts, closing
breakage and injection holes. - Each launch picks a unique rendezvous port instead of a hardcoded one.
- Ephemeral blocking launches stream output, are supervised, and expand
their CLI environment in-process; non-blocking immutable scripts now
receive the full launcher environment.
Scheduler backends and system profiles
- LSF: bsub flags are proper argv tokens, submit-only flags no longer
leak into jsrun, get_job_id is implemented, and --smpiargs quoting is
fixed. Sierra CPU-affinity binding follows the requested procs/node,
and the rzansel hostname typo means RZAnsel is detected again. - Scheduler argument dicts and System env lists are per-instance, so
state no longer leaks between jobs in one process; SystemParams is
copied before overrides; scheduler autodetection and -p scheduler=
are honored. - The RCCL/NCCL environment derives from torch's bundled ROCm rather
than ROCM_PATH, NCCL_NET is set only when the libfabric plugin is
actually present, unversioned ROCM_PATH no longer crashes, and El
Capitan MIOpen cache paths tolerate an unset TMPDIR. Adds a Corona
profile and a ROCm 7.x PyTorch behavior variable.
CLI and packaging
- Torchrun passthrough arguments respect the launcher/script boundary,
handle abbreviated flags, and validate identically across rendezvous
protocols; module execution via -m is supported; -x/--xargs accepts
the schedulers' real dashed keys; missing commands and same-file
--out/--err fail with clean errors; resource-validation failures are
visible and actionable. - GPU vendor libraries moved from install_requires to [rocm]/[cuda]
extras, so wheels no longer bake in the build machine's hardware; a
new [rocm-auto] extra pins amdsmi to the installing machine's ROCm
release for source installs. - Launch directory names include a UUID for uniqueness; documentation
that contradicted actual behavior was corrected.
Tests and CI
- Extensive new regression coverage (GPU visibility, argument
boundaries, state leaks, packaging reproducibility, docs consistency);
several tautological or shadowed tests now assert what they meant to;
torch is installed on every CI leg and its absence fails instead of
skipping; the minimum Python for CI was raised.
v1.0.4
What's Changed
- Improve the performance for ROCm 7 + Slingshot systems to find and use the AWS_OFI_NCCL plugin. This became the default plugin and deprecated the use of the AWS_OFI_RCCL plugin (which is still used for ROCm 6.x)
- Added environment variables to force use of the
libfabricinterface provided by the AWS_OFI_*CCL plugin's and is required for high performance communication. Can be overridden by explicitly setting theNCCL_NETenvironment variable. - Added performance tuning flags for ROCm systems in PyTorch to use channel's last ordering, which aligns with the best practices for the ROCm / MIOpen libraries.
Full Changelog: v1.0.3...v1.0.4
v1.0.3
v1.0.3 release:
Setup scripts to autorelease to PyPI
Update systems and bug fix autodetect (#57)
-
Fixed a bug in the autodetect GPU logic that overwrote the output in
the finally block attached to the try block. Added code to
auto-detect the ROCm version being used and set that to constrain the
version of amdsmi installed. Added definitions for more LLNL systems. -
Added installation instructions.
Added the device_id initialization to the init_process_group call in
the torchrun-hpc trampoline.
v1.0.2
v1.0.1
What's Changed
- Fixed license to use SPDX format. by @bvanessen in #53
Full Changelog: v1.0.0...v1.0.1
v1.0.0
The HPC launcher repository contains a set of helpful scripts and Python bindings for launching PyTorch (torchrun), LBANN 2.0 (PyTorch-core), or generic scripts on multiple leadership-class HPC systems. There are optimized routines for FLUX, SLURM, and LSF launchers. Additionally, there are optimized environments for systems at known compute centers. Currently there are supported systems at:
- LLNL Livermore Computing (LC)
There are two main entry points into HPC-Launcher from the cli: launch and torchrun-hpc. torchrun-hpc is intended as a replacement for torchrun, while launch is a generic interface for launching parallel jobs.