·
114 commits
to main
since this release
Immutable
release. Only release title and notes can be modified.
NVIDIA Resiliency Extension v0.7.0
Highlights
-
In-job restart and resilient deployment
- Barrier rendezvous is now the only FT rendezvous implementation (#379). The compatibility spelling
--ft-rdzv-impl barrierremains accepted;legacyis rejected. - Hot-spare restart rounds prefer returning active nodes over early standbys and exclude every participant from an unhealthy Slurm array task (#325, #326). Recovery is also hardened against transient over-subscription and stale participants from failed replacement groups (#357, #391).
- Scheduler segment-health integration lets each allocation unit publish and consume shared exclusion decisions before rendezvous, so unhealthy Slurm array tasks are omitted when eligible spares are available (#398, #399). See the Segment Health Check guide.
- A new singleton job-array deployment example combines in-job hot spares with scheduler-managed replacement of an unrecoverable rendezvous-host allocation (#383). See the deployment guide.
- Successful node-health responses with unusable output now fail open; only an explicit service failure or a numeric
fail_count > 0marks the node unhealthy (#420). - Local launcher failures no longer permanently close the shared rendezvous. Terminal worker-group failures and genuine job-wide shutdown decisions continue to close it (#350, #421).
- Progress-tracker values supplied through YAML are preserved when the corresponding CLI options are omitted (#329).
- Barrier rendezvous is now the only FT rendezvous implementation (#379). The compatibility spelling
-
Attribution and restart decisions (experimental)
ft_launchercan run or contact the Attribution Service and submit per-cycle logs for restart analysis (#338).- Attribution analysis is asynchronous and does not delay restart. Acting on a
STOPverdict is opt-in through--ft-attribution-stop-action no-restart; the defaultlogmode records verdicts without terminating the job (#378). - The new Restart Agent provides deterministic failure evidence and retry policy, with optional model-assisted enrichment, bounded history, operational logging, and auditable result artifacts (#370, #400).
- The core wheel includes
attrsvc, Flight Recorder analysis, and the direct Restart Agent backend. MCP and Slack integrations requirenvidia-resiliency-ext[attribution]; legacy LogSage remains source-checkout-only and is excluded from wheels (#404, #410).
-
Health checks
NVLinkWindowHealthCheckdetects persistent NVLink replay and recovery counter growth across a configurable sampling window while ignoring isolated transient changes (#302).- The default
nvhcdnode-health request is limited to theprolog,epilog,logs, andgpuDCAHC groups; explicit caller arguments still override that default (#356).
-
Checkpointing
-
Logging, packaging, and security
- gRPC log servers confine client-selected write paths to launcher-derived allowed roots and reject traversal, symlink escape, and other out-of-root destinations (#415).
- Multi-node direct append to a shared per-cycle log is no longer recommended. Enable the gRPC log funnel for a single shared-filesystem writer; direct append can serialize on Lustre and can lose or interleave data on NFS-style storage (#418).
- Release wheels and CI target Python 3.12 and 3.14, including an
ft_launchercompatibility fix for Python 3.14. Python 3.10 and 3.11 are no longer covered by the wheel/unit-test matrix (#342, #372). - The optional MCP SDK is constrained to
>=1.28.1,<2.0.0, incorporating the fix for CVE-2026-59950 while retaining compatibility with the NVRx MCP integration (#388, #412).
Deprecations & Removals
- Legacy dynamic rendezvous has been removed. Remove
--ft-rdzv-impl legacy; omit the option or use--ft-rdzv-impl barrier(#379). --ft-restart-policyremains deprecated. Onlyany-failedis supported and it is already the default, so remove the option from launch scripts.- In-process restart remains deprecated. New deployments should use in-job restart through
ft_launcher. - The
[dataflow]extra andnvdataflowdependency have been removed. Attribution export uses the configured direct HTTP endpoint instead (9182c06).
Installation
# Core resiliency, attrsvc, and the direct Restart Agent backend
pip install nvidia-resiliency-ext
# Optional MCP and Slack attribution integrations
pip install 'nvidia-resiliency-ext[attribution]'Known Issues & Limitations
- Linux wheel compatibility: published wheels target
manylinux_2_39; systems with glibc older than 2.39, including standard Ubuntu 22.04 installations, must build from source. - Python: release wheels are published for Python 3.12 and 3.14. The package metadata still permits Python 3.10+, but Python 3.10 and 3.11 are not covered by the release wheel/CI matrix.
- Attribution and Restart Agent features are experimental. APIs, service contracts, schemas, and model-route configuration may change.
- External attribution endpoints: the in-job client currently submits logs over HTTP(S). Other endpoint schemes do not provide an additional transport implementation.
- Segment-health decisions fail open when their artifact is missing or unusable. Slurm does not evict an already running allocation merely because a node becomes drained; the decision is consumed at the next rendezvous health-check boundary.
- Per-cycle log aggregation is best-effort. Failure-adjacent output may be incomplete; retain launcher/rank-monitor logs and core dumps for authoritative diagnostics.