TensorFold 0.3.5.1
A fix for M1 and M2 Macs: Qwen3.8-27B wouldn't load on them. 0.3.4.1 stopped with "Thread group size (1024) is greater than the maximum allowed threads per threadgroup (704)", and 0.3.5 with 512 over 448.
Metal gives each compiled kernel its own limit on threads per threadgroup, and on M1 and M2 that limit falls as the kernel uses more registers. M3 and later give every kernel 1024, which is why our test machines never hit it.
- On M1 and M2 the 4-bit row matmul checks each of its kernels the first time it runs and uses fewer simdgroups where 16 don't fit. The sums run in the same order, so drafted output still equals serial output.
- Every other kernel above 256 threads now declares its size to the Metal compiler, which makes M1 and M2 fit it: the norms, sampling and top-k, Nemotron's norms and row matmul, and Flash Next's larger kernels.
- On M3, M4 and M5 nothing changes: the compiled kernels are 0.3.5's machine code. Paired against 0.3.5 on an M3 Ultra, every token hash matches, and decode and cold prefill from 2k to 32k tokens stay level with 0.3.5 (0.99-1.02x).
Thanks to @hichaiuse, @simonmd, @gcarusso, @tonydehnke, @Cyb3r-Monk and @tinyapps for the reports, the repro and the numbers that found it.
Install
tensorfold update installs this release. Or:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.5.1