Skip to content

b11372

Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 03 Oct 11:24
889edf4

qwen4exp : halve the indexer score memory (#29825)

  • qwen4exp : halve the indexer score memory

The indexer scored all heads in one product and rectified a copy of it,
so two [n_pool, n_idx_h, n_tokens] f32 tensors were live at once, the
largest buffers of the graph at long context. Each head now gets its
own product, rectified and summed in place into one [n_pool, n_tokens]
score.

  • qwen4exp: let the allocator reuse the indexer score buffers

Address review from CISC: use plain ggml_add and ggml_relu in the
indexer head loop. The graph allocator already runs them in place when
their source has no other consumer, so the _inplace variants are not
needed. The compute buffer and the speed are unchanged.

  • cuda: support 4 heads in the lightning indexer

Dispatch 4 heads to the vector kernel, too few for a wmma tile, and
accept them in supports_op. test-backend-ops covers 4 heads.

  • metal: take the lightning indexer head count as a function constant

The kernel reads the head count from a function constant and zero fills
the last head tile, so any head count runs and 64 heads is unchanged.

  • qwen4exp: compute the indexer score with the lightning indexer

Address review from am17an: the unweighted sum of the rectified head
scores scaled by 1/sqrt(head_dim) is the lightning indexer with every
head weight set to that scale, so the indexer calls
ggml_lightning_indexer on the pooled keys with an f16 pool mask. The
keys are read once for all heads and no per head score is
materialized.

  • vulkan: tile the lightning indexer over keys and tokens

A workgroup scores 64 keys against 8 tokens: the keys are staged once
in shared memory, the queries one head at a time, and each invocation
owns one key for two tokens, so no dot product needs a cross invocation
reduction. The subgroup variant and the flat dispatch are gone, the grid
is keys x tokens x streams.

  • vectorize vulkan loads and use fp16 dot product

Co-authored-by: Ruben Ortlam rortlam@redhat.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI: