Skip to content

b11482

Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 08 Oct 02:24
b86d2f0

cuda: FWHT kernels for block widths above 512 (#29100)

The CUDA FWHT covers widths 64 to 512. It runs one row per warp and keeps N/32
values per lane, so wider blocks need more registers per lane than that layout
allows.

fwht_cuda_block runs one row per thread block with 256 threads, so each thread
keeps N/256 values. Stages below the warp width still shuffle, those up to the
block width go through shared memory, and the rest stay in registers. Same
butterfly and sign convention as the warp kernel.

Widths 64 to 512 keep the warp kernel. 1024 through 8192 use the new one, for
both F32 and F16 sources. ggml_cuda_op_mul_mat_use_fwht (the shared
supports_op/dispatch predicate added in #29096) does not check width, so it
needed no change here: any width it admits that ggml_cuda_op_fwht can't serve
already falls through correctly to the cuBLAS path.

Rebased onto current master with #29096's F16 commit underneath it, since this
depends on the same F16 template infrastructure; that commit applied cleanly,
the only conflict was in test-backend-ops.cpp where an unrelated intervening
commit's own test additions landed near this block.

test-backend-ops on an A10 (lambdalabs): MUL_MAT 1297/1297, including all
FWHT/Hadamard cases (18 existing, 4 new F32 wide, 4 new F16 wide, 2 new
many-rows, 1 too-big boundary moved to 16384).

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI: