[ExecuTorch][WebGPU] Replace broad QKV fusion with BK64 kernel#21131
[ExecuTorch][WebGPU] Replace broad QKV fusion with BK64 kernel#21131JCNTH wants to merge 4 commits into
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21131
Note: Links to docs will display an error until the docs builds have been completed. ❗ 1 Active SEVsThere are 1 currently active SEVs. If your PR is affected, please view them below: ❌ 2 New Failures, 1 Unrelated FailureAs of commit 9d1f0c1 with merge base 86c3470 ( NEW FAILURES - The following jobs have failed:
BROKEN TRUNK - The following job failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
Stack from ghstack (oldest at bottom):
The previously landed broad QKV fusion applied too widely and did not match the
BK64 schedule now used for the ordinary projections. This corrective diff
replaces it with a capability- and geometry-qualified BK64 kernel that fuses the
exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the
constant weights and scales once and scattering the result into three distinct
planner-safe outputs. Outside the accepted shapes it switches atomically back
to the ordinary Steel and bicol routes, and it deletes the obsolete broad
shader and header so only one QKV path remains. Mirrors Vulkan
xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl
for the per-projection GEMM; the three-output fusion itself is WebGPU-specific.
Key changes:
removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header.
packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol.
@exported-using-ghexport
Differential Revision: D113171749
Differential Revision: D113171749