TensorFold 0.3.6.1: CUDA builds inside NVIDIA's containers
A fix for CUDA builds inside NVIDIA's containers. They set TORCH_CUDA_ARCH_LIST to every architecture back to sm_80, so 0.3.5 to 0.3.6 compiled the kernels' thread-block clusters and FP8 MMA for GPUs without them and stopped with namespace "cooperative_groups" has no member "this_cluster".
- Every CUDA extension now builds for the GPU that is present, so the container's list adds nothing.
- A GPU older than compute capability 9.0 is refused with a message that names it.
- Checked on a DGX Spark under the container's full list with a fresh extension cache: every extension built for sm_121, and the CUDA suite passed (433 tests). The Mac suite passes on an M5 Max (1,779 tests).
Thanks to @ss-cong for the report and the exact errors (#56).
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.6.1