[ExecuTorch][WebGPU] Coordinate dynamic dispatch routes#21128
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21128
Note: Links to docs will display an error until the docs builds have been completed. ❗ 1 Active SEVsThere are 1 currently active SEVs. If your PR is affected, please view them below: ❌ 2 New Failures, 1 Unrelated FailureAs of commit fabbff4 with merge base 86c3470 ( NEW FAILURES - The following jobs have failed:
BROKEN TRUNK - The following job failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
Stack from ghstack (oldest at bottom):
Quantized-linear and attention optimizations each want to swap in a specialized
dispatch range at runtime depending on the live sequence length, but nothing
prevented two routes from claiming the same dispatch slot or from running an
inactive zero-grid route. This adds a central registry of mutually exclusive
compute-dispatch ranges and selects the quantized-linear and SDPA route for
each execution from the live dynamic shape without rebuilding the graph.
Inactive zero-grid routes are excluded from execution and timestamp indexing;
invalid, overlapping, or copy-containing ranges are rejected; static or
ineligible graphs keep the established lowering. This is shared infrastructure
the GEMM and attention routes build on. The mutually-exclusive route registry
itself is WebGPU-specific; it builds on the same per-execution dynamic-resize
pattern as Vulkan's DynamicDispatchNode/ExecuteNode resize hooks.
Key changes:
dynamic shapes.
ranges through the registry.
@exported-using-ghexport
Differential Revision: D113171748
Differential Revision: D113171748