RFC: HRX (Loom) + NPU, and Vulkan out once HRX meets the gates #213
Unanswered
bong-water-water-bong
asked this question in
rfcs
Replies: 1 comment
|
Credit where it's due: this direction is geramyL's (a moderator on Lemonade's Discord). Reviewing the Laya work, they said "we want it on the NPU / vulkan version of it already exists / HRX + LOOM + NPU". We moved only the Laya scorer to HRX then (#208). The owner now agrees it's the right call for the whole engine, and this RFC is how we get there without losing speed or models on the way. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Proposal. The engine becomes HRX (AMD's ggml-hrx, kernels written in Loom) for the GPU plus the NPU, and Vulkan is removed. The removal lands only once HRX meets the gates below; until then Vulkan stays the default and nothing is taken out.
Why. HRX is AMD's own backend for this hardware, its kernels are ours to write and fix (the #123/#140 race fix, the ZAYA and Laya kernels, #180), and it is the path the NPU and GPU work already share. Carrying two GPU stacks doubles the pins (
third_party/llama.cpp-vulkanand the HRX build), the bump workflows, the packaging and the test matrix.What Vulkan does today (main at f362a09).
config/route-policy.jsonsends every Laya class tovulkan, and--device autopicks it for GGUFs.--mtp), DFlash2 (--dflash), and Zamba2 (carried ggml-org#21412). Qwen3.8-Flash-Next loads on Vulkan and fails to load on the HRX pin.Where HRX stands. It maps nearly the same models: 94.77% of HF text-generation models against Vulkan's 94.87%, and 64.31% against 64.32% checked end to end (census 2026-09-28). Its prompt processing already beats Vulkan's on long prompts (HRX prefill + Vulkan decode: -26% time on an 8K prompt, Qwen2.5-7B). Known issues (docs/hrx.md): several sequences per batch fail (3-D MUL_MAT),
-fa offfails, and each new shape pays a JIT compile on first use.Gates before removal (each measured on Strix Halo, numbers in the PR that removes Vulkan):
--mtp,--dflash,--parallel(several sequences per batch) and--mmprojwork on HRX.serve_e2e,smoke_serveand the Lemonade recipe test pass with no Vulkan in the build.Then, in one PR: drop
third_party/llama.cpp-vulkanandbump-llama-vulkan.yml, thevulkandevice in1bit serveandroute-policy.json, the Vulkan builds in packaging,weekly.shandbuild-windows.sh, and the Vulkan columns in the registry; move PR-Agent to HRX; update the docs, the site and NOTICE.Not in scope. The ROCm lean route (Hadamard W4A4), ZINC and DwarfStar are separate backends; each gets its own decision.
Open questions. Windows builds use Vulkan (
build-windows.sh); HRX has no Windows path yet. The ONNX WebGPU provider (Dawn) sits on Vulkan too.All reactions