Skip to content

b10840

Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 07 Sep 14:31
73ab759

CUDA: branchless Q4_K/Q5_K unpack to speed up mmvq, L2 prefetch on DGX Spark (#26705)

  • Update Q4_K and Q5_K to use branchless computation, which stops the scale unpack being re-executed for every column in mmvq, improving perf at batch sizes > 1

  • Gating the change off from DGX Spark due to no gain

  • Adding prefetch gated to Spark, making branchless change in Q4_K and Q5_K general and modifying switch points based on latest perf data

  • Guard the mmvq L2 prefetch against MUSA as well as HIP

  • Define the mmvq L2 prefetch only under the Spark guard

  • Update switch point for Q4_K to accommodate more models

  • Remove stale comments

  • Add block_size to ggml_cuda_type_traits and create a separate mmvq_should_prefetch function

  • Rename block_size to bs for cleaner indentation

  • Fix build error on non-Spark CUDA arch with appropriate conditional around new function added


Co-authored-by: praneshgo 227579474+praneshgo@users.noreply.github.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI: