Skip to content

TensorFold 0.3.6.1: CUDA builds inside NVIDIA's containers

Choose a tag to compare

@ashhart ashhart released this 28 Sep 09:37
· 32 commits to main since this release

A fix for CUDA builds inside NVIDIA's containers. They set TORCH_CUDA_ARCH_LIST to every architecture back to sm_80, so 0.3.5 to 0.3.6 compiled the kernels' thread-block clusters and FP8 MMA for GPUs without them and stopped with namespace "cooperative_groups" has no member "this_cluster".

  • Every CUDA extension now builds for the GPU that is present, so the container's list adds nothing.
  • A GPU older than compute capability 9.0 is refused with a message that names it.
  • Checked on a DGX Spark under the container's full list with a fresh extension cache: every extension built for sm_121, and the CUDA suite passed (433 tests). The Mac suite passes on an M5 Max (1,779 tests).

Thanks to @ss-cong for the report and the exact errors (#56).

Install

tensorfold update installs this release. Or:

pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6.1