Skip to content

v0.7.0

Latest

Choose a tag to compare

@hexinw-nvidia hexinw-nvidia released this 16 Sep 01:09
· 114 commits to main since this release
Immutable release. Only release title and notes can be modified.
ebf0352

NVIDIA Resiliency Extension v0.7.0

Highlights

  • In-job restart and resilient deployment

    • Barrier rendezvous is now the only FT rendezvous implementation (#379). The compatibility spelling --ft-rdzv-impl barrier remains accepted; legacy is rejected.
    • Hot-spare restart rounds prefer returning active nodes over early standbys and exclude every participant from an unhealthy Slurm array task (#325, #326). Recovery is also hardened against transient over-subscription and stale participants from failed replacement groups (#357, #391).
    • Scheduler segment-health integration lets each allocation unit publish and consume shared exclusion decisions before rendezvous, so unhealthy Slurm array tasks are omitted when eligible spares are available (#398, #399). See the Segment Health Check guide.
    • A new singleton job-array deployment example combines in-job hot spares with scheduler-managed replacement of an unrecoverable rendezvous-host allocation (#383). See the deployment guide.
    • Successful node-health responses with unusable output now fail open; only an explicit service failure or a numeric fail_count > 0 marks the node unhealthy (#420).
    • Local launcher failures no longer permanently close the shared rendezvous. Terminal worker-group failures and genuine job-wide shutdown decisions continue to close it (#350, #421).
    • Progress-tracker values supplied through YAML are preserved when the corresponding CLI options are omitted (#329).
  • Attribution and restart decisions (experimental)

    • ft_launcher can run or contact the Attribution Service and submit per-cycle logs for restart analysis (#338).
    • Attribution analysis is asynchronous and does not delay restart. Acting on a STOP verdict is opt-in through --ft-attribution-stop-action no-restart; the default log mode records verdicts without terminating the job (#378).
    • The new Restart Agent provides deterministic failure evidence and retry policy, with optional model-assisted enrichment, bounded history, operational logging, and auditable result artifacts (#370, #400).
    • The core wheel includes attrsvc, Flight Recorder analysis, and the direct Restart Agent backend. MCP and Slack integrations require nvidia-resiliency-ext[attribution]; legacy LogSage remains source-checkout-only and is excluded from wheels (#404, #410).
  • Health checks

    • NVLinkWindowHealthCheck detects persistent NVLink replay and recovery counter growth across a configurable sampling window while ignoring isolated transient changes (#302).
    • The default nvhcd node-health request is limited to the prolog, epilog, logs, and gpu DCAHC groups; explicit caller arguments still override that default (#356).
  • Checkpointing

    • Persistent async checkpoint workers refresh cached CUDA IPC handles when tensor storage changes, preventing stale-buffer reuse (#314).
    • CPU tensors are cloned before persistent async serialization so optimizer steps and other live CPU state cannot mutate an in-flight checkpoint (#376).
  • Logging, packaging, and security

    • gRPC log servers confine client-selected write paths to launcher-derived allowed roots and reject traversal, symlink escape, and other out-of-root destinations (#415).
    • Multi-node direct append to a shared per-cycle log is no longer recommended. Enable the gRPC log funnel for a single shared-filesystem writer; direct append can serialize on Lustre and can lose or interleave data on NFS-style storage (#418).
    • Release wheels and CI target Python 3.12 and 3.14, including an ft_launcher compatibility fix for Python 3.14. Python 3.10 and 3.11 are no longer covered by the wheel/unit-test matrix (#342, #372).
    • The optional MCP SDK is constrained to >=1.28.1,<2.0.0, incorporating the fix for CVE-2026-59950 while retaining compatibility with the NVRx MCP integration (#388, #412).

Deprecations & Removals

  • Legacy dynamic rendezvous has been removed. Remove --ft-rdzv-impl legacy; omit the option or use --ft-rdzv-impl barrier (#379).
  • --ft-restart-policy remains deprecated. Only any-failed is supported and it is already the default, so remove the option from launch scripts.
  • In-process restart remains deprecated. New deployments should use in-job restart through ft_launcher.
  • The [dataflow] extra and nvdataflow dependency have been removed. Attribution export uses the configured direct HTTP endpoint instead (9182c06).

Installation

# Core resiliency, attrsvc, and the direct Restart Agent backend
pip install nvidia-resiliency-ext

# Optional MCP and Slack attribution integrations
pip install 'nvidia-resiliency-ext[attribution]'

Known Issues & Limitations

  • Linux wheel compatibility: published wheels target manylinux_2_39; systems with glibc older than 2.39, including standard Ubuntu 22.04 installations, must build from source.
  • Python: release wheels are published for Python 3.12 and 3.14. The package metadata still permits Python 3.10+, but Python 3.10 and 3.11 are not covered by the release wheel/CI matrix.
  • Attribution and Restart Agent features are experimental. APIs, service contracts, schemas, and model-route configuration may change.
  • External attribution endpoints: the in-job client currently submits logs over HTTP(S). Other endpoint schemes do not provide an additional transport implementation.
  • Segment-health decisions fail open when their artifact is missing or unusable. Slurm does not evict an already running allocation merely because a node becomes drained; the decision is consumed at the next rendezvous health-check boundary.
  • Per-cycle log aggregation is best-effort. Failure-adjacent output may be incomplete; retain launcher/rank-monitor logs and core dumps for authoritative diagnostics.