Skip to content

Releases: llnl/HPC-launcher

v1.0.6

Choose a tag to compare

@bvanessen bvanessen released this 15 Sep 05:58

HPC-launcher v1.0.6

Nested job steps inside existing allocations (#71)

Blocking launches invoked from inside a scheduler allocation (salloc, sbatch,
flux alloc) now run as a job step in the current allocation instead of
accidentally requesting a new one.

  • Inherit allocation size: with no job-size flag (-N/-g/--gpumem-at-least),
    the node count is inherited from the enclosing allocation via a new
    scheduler-agnostic probe covering Flux, SLURM, and LSF. Explicit flags still
    take precedence.
  • New -c/--cpus-per-task flag to give a job or step an explicit CPU
    footprint (SLURM --cpus-per-task, Flux --cores-per-task/--cores-per-slot).
  • SLURM job steps: allocation-selection flags (--partition, --account,
    --reservation) are dropped with a warning inside an allocation; node requests
    exceeding the allocation are rejected upfront; and --overlap is added by
    default so concurrent steps don't hang on busy nodes (omitted with
    --cpus-per-task or --exclusive, removable with -x ~--overlap).
  • Non-blocking (sbatch/flux batch) submissions are unchanged and still create
    a new job.

Includes 17 new scheduler-free tests (tests/nested_job_step_test.py) and
updated CLI documentation.

v1.0.5

Choose a tag to compare

@bvanessen bvanessen released this 03 Aug 04:29
dbbb021

Release v1.0.5

A large bug-fix and hardening release driven by two rounds of deep code
review, with end-to-end validation on El Capitan-class (MI300A/Flux)
systems.

GPU assignment and the torchrun trampoline

  • Workers no longer initialize HIP/CUDA at import time with all GPUs
    visible, which pinned every rank on a node to physical GPU 0; each
    rank now sees only its granted devices, and the memory-fraction cap
    applies to the device actually selected.
  • The trampoline puts the user's own code on sys.path (sibling imports
    work under --local, subdirectories, and custom launch dirs), destroys
    the process group on every exit path, publishes rank identity itself,
    and only passes device_id to init_process_group for accelerators, so
    CPU/gloo jobs start again.

Correctness of generated launches

  • Job exit codes propagate through Scheduler.launch to the CLI exit
    status; SIGINT kills the scheduler child instead of orphaning it.
  • User-controlled values are shell-quoted in generated scripts, closing
    breakage and injection holes.
  • Each launch picks a unique rendezvous port instead of a hardcoded one.
  • Ephemeral blocking launches stream output, are supervised, and expand
    their CLI environment in-process; non-blocking immutable scripts now
    receive the full launcher environment.

Scheduler backends and system profiles

  • LSF: bsub flags are proper argv tokens, submit-only flags no longer
    leak into jsrun, get_job_id is implemented, and --smpiargs quoting is
    fixed. Sierra CPU-affinity binding follows the requested procs/node,
    and the rzansel hostname typo means RZAnsel is detected again.
  • Scheduler argument dicts and System env lists are per-instance, so
    state no longer leaks between jobs in one process; SystemParams is
    copied before overrides; scheduler autodetection and -p scheduler=
    are honored.
  • The RCCL/NCCL environment derives from torch's bundled ROCm rather
    than ROCM_PATH, NCCL_NET is set only when the libfabric plugin is
    actually present, unversioned ROCM_PATH no longer crashes, and El
    Capitan MIOpen cache paths tolerate an unset TMPDIR. Adds a Corona
    profile and a ROCm 7.x PyTorch behavior variable.

CLI and packaging

  • Torchrun passthrough arguments respect the launcher/script boundary,
    handle abbreviated flags, and validate identically across rendezvous
    protocols; module execution via -m is supported; -x/--xargs accepts
    the schedulers' real dashed keys; missing commands and same-file
    --out/--err fail with clean errors; resource-validation failures are
    visible and actionable.
  • GPU vendor libraries moved from install_requires to [rocm]/[cuda]
    extras, so wheels no longer bake in the build machine's hardware; a
    new [rocm-auto] extra pins amdsmi to the installing machine's ROCm
    release for source installs.
  • Launch directory names include a UUID for uniqueness; documentation
    that contradicted actual behavior was corrected.

Tests and CI

  • Extensive new regression coverage (GPU visibility, argument
    boundaries, state leaks, packaging reproducibility, docs consistency);
    several tautological or shadowed tests now assert what they meant to;
    torch is installed on every CI leg and its absence fails instead of
    skipping; the minimum Python for CI was raised.

v1.0.4

Choose a tag to compare

@bvanessen bvanessen released this 22 Jan 01:57
0f65e74

What's Changed

  • Improve the performance for ROCm 7 + Slingshot systems to find and use the AWS_OFI_NCCL plugin. This became the default plugin and deprecated the use of the AWS_OFI_RCCL plugin (which is still used for ROCm 6.x)
  • Added environment variables to force use of the libfabric interface provided by the AWS_OFI_*CCL plugin's and is required for high performance communication. Can be overridden by explicitly setting the NCCL_NET environment variable.
  • Added performance tuning flags for ROCm systems in PyTorch to use channel's last ordering, which aligns with the best practices for the ROCm / MIOpen libraries.

Full Changelog: v1.0.3...v1.0.4

v1.0.3

Choose a tag to compare

@bvanessen bvanessen released this 08 Jan 21:39
4c97b24

v1.0.3 release:

Setup scripts to autorelease to PyPI

Update systems and bug fix autodetect (#57)

  • Fixed a bug in the autodetect GPU logic that overwrote the output in
    the finally block attached to the try block. Added code to
    auto-detect the ROCm version being used and set that to constrain the
    version of amdsmi installed. Added definitions for more LLNL systems.

  • Added installation instructions.

Added the device_id initialization to the init_process_group call in
the torchrun-hpc trampoline.

v1.0.2

Choose a tag to compare

@bvanessen bvanessen released this 31 Aug 11:59
b44575d

What's Changed

Full Changelog: v1.0.1...v1.0.2

v1.0.1

Choose a tag to compare

@bvanessen bvanessen released this 31 Aug 11:50
5ca955c

What's Changed

Full Changelog: v1.0.0...v1.0.1

v1.0.0

Choose a tag to compare

@bvanessen bvanessen released this 27 Aug 17:56
b479c41

The HPC launcher repository contains a set of helpful scripts and Python bindings for launching PyTorch (torchrun), LBANN 2.0 (PyTorch-core), or generic scripts on multiple leadership-class HPC systems. There are optimized routines for FLUX, SLURM, and LSF launchers. Additionally, there are optimized environments for systems at known compute centers. Currently there are supported systems at:

  • LLNL Livermore Computing (LC)

There are two main entry points into HPC-Launcher from the cli: launch and torchrun-hpc. torchrun-hpc is intended as a replacement for torchrun, while launch is a generic interface for launching parallel jobs.