Skip to content

NCCL v2.32.3-1 Release

Latest

Choose a tag to compare

@bhramesh-nvidia bhramesh-nvidia released this 17 Sep 21:09
12df1a1

Vera Rubin Support

  • Adds initial Rubin platform support including support for sm107, CX9 rail and plane detection, and MPS+MLoPart.
  • 2.32.3 focuses on new functionalities and does not contain performance model tuning for Rubin. This will be part of the next release to improve out-of-box experience for Rubin users.

Device API Enhancements

  • Adds Compute Fabric Transport (CFT) counted-write and wait support.
  • Adds socket-based GIN support, enabling custom kernel development with TCP sockets.
  • Adds support in GDAKI for LAG-aware QP assignment based on the context ID. (Github PR #2315)
  • Adds NCCL_WIN_GIN_ONLY so users can register a window only for GIN usage.
  • Optimize GIN performance by skipping mcst operation when possible.

Collectives and Runtime Enhancements

  • Adds a ring-based hierarchical copy-engine AllGather implementation selectable with NCCL_HIER_CE_COLL_AG_RAIL_RING_ENABLE. (Github PR #2299)
  • Improves Blackwell symmetric AllGather performance, resource overhead modeling and kernel selection with a new cost model.
  • Adds optional TLS encryption for NCCL-owned socket traffic when NCCL is built with OpenSSL3, configured through the ncclSetEncryption API.
  • Adds ncclCollConfig_t::launchCompletionEvent, allowing callers to observe kernel-launch completion.
  • Adds ncclNvlsHostMode_t to ncclConfig_t so applications can disable host NVLS collectives per communicator.
  • Optimizes NVLS slot consumption to avoid resource exhaustion issues when using multiple communicators with NVLS.

Diagnostics and Profiling

  • Adds the ATTN log level for important non-fatal conditions, such as configuration fallbacks and plugin initialization failures.
  • Expands RAS capability with GPU-resident progress counters and a watchdog DMA mirror for diagnosing stalled collective kernels.
  • Adds additional RAS diagnostics functionalities for NVLink/NIC state and speed, PCI and GDR configuration, Xid/SXid events and others.
  • Reports degraded NVLink fabric bandwidth with an ATTN message during communicator initialization.
  • Reports mismatched NCCL Git revisions across communicator ranks during initialization.

Other Improvements

  • Fixes PAT+NVLS performance drops on H100 platforms.
  • Avoid address space exhaustion when repeatedly registering symmetric windows backed by a single physical memory allocation.
  • Reduces communicator initialization overhead at large scale by avoiding scans of inactive proxy poll descriptors.
  • Adds an experimental built-in NetworkDirect transport on Windows, with automatic socket fallback.
  • EFA team improved EFA GDA support with gin.get API and others. (Github PR #2382)(Github PR #2398)(Github PR #2405)
  • Reports Inspector ring-buffer drops and operation counts in JSON. (Github PR #2304)
  • NCCL now supports communicators using multiple MIG instances.
  • Generates llms.txt for agents to better read NCCL documentation.

Bug Fixes

  • Fixes profiler overhead when no profiler plugin is loaded. (Github Issue #2355)
  • Fixes P2P IPC registration reuse producing out-of-bounds remote addresses when a registered allocation spans multiple cuMem segments. (Github PR #2362)
  • Fixes PAT connection-setup deadlocks when runtime connection is disabled and nodes have uneven local-rank counts. (Github Issue #2385)
  • Fixes non-thread-safe token parsing that could corrupt concurrent configuration parsing. (Github Issue #2361)
  • Fixes B40 TMA symmetric kernels crash due to insufficient shared memory.
  • Fix GIN GDAKI bug related to hop limit when using DOCA SDK.
  • Fixes GIN Proxy and GPI descriptor shared-memory sizing and alignment.
  • Fixed CFT window registration performed before ncclDevCommCreate.
  • Fixes profiler API events reporting rank 0 instead of the originating communicator rank. (Github Issue #2300)
  • Fixes tuner plugins receiving uninitialized cost-model constants after the tuning rework.

Contrib/ Updates

  • Adds more functionality in NiiN.
  • Fixes contrib/nccl_checkpoint rejecting otherwise compatible NCCL versions. (Github Issue #2347)

Acknowledgements

We thank the following contributors for their work on this release:

@madeleineth, @alpha-baby, @akkart-aws, @anshumang, @cesar-stuardo-bd, @hexagonal-banana,
@rauteric, @shaq918, @yshalabi, @zrss for your contributions.

We also thank the community for issue reports, testing, and feedback.

Known Issues

  • Rubin performance model has not been optimized with 2.32.3. Users can observe better performance on Rubin by increasing the number of CTAs NCCL uses. Automatic tuning will be improved in the next release.
  • Socket GIN currently requires users to opt-in to enable GDRCopy.
  • B100 PCIe MLoPart workloads can fail during memory allocation with CUDA error 101 (invalid device ordinal). Set NCCL_CUMEM_ENABLE=0 as a workaround.