Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines over NVSHMEM symmetric memory. Zero SM usage; 1.7-2.2x over torch.distributed on NVLink.
cuda ulysses pytorch diffusion-models collective-communication nvshmem sequence-parallelism custom-operator
-
Updated
Aug 7, 2026 - Python