v0.8.17 — trellis MoE serves on CPU
A trellis MoE checkpoint now serves on a machine with no GPU. Until 0.8.16 the vLLM
CPU backend refused every GLQ MoE, which put the most CPU-suited architecture we publish
out of reach: an MoE reads only its top-k experts per token, so it moves far less memory
than a dense model of the same size.
What changed
- CPU MoE serving. The refusal is per codebook now — trellis is allowed, e8p/shell
still refuse at load with the reason and the fix, because their expert decode is
CUDA-only. - The fused CPU expert kernel runs under vLLM.
GLQFusedMoEMethod.apply()takes a CPU
branch: one extension call per MoE block instead of a Python loop of per-expert calls.
The CUDA path is untouched. Anything outside the op's reach (stage-2 RVQ, padded shapes,
ungated activations, an older wheel) falls back to the loop — correct on any shape — and
says why, once, rather than dropping onto the slower path silently. - The installer recommends it. A trellis MoE that fits the RAM budget is now the CPU
recommendation; e8p/shell and unknown formats are still excluded, and the budget is
unchanged. glq-chatclaims the core vLLM holds back (VLLM_CPU_NUM_OF_RESERVED_CPU=0),
worth ~11% on a 4-core box. Set it yourself to keep a core reserved.
Breaking: Python 3.12 is now the floor
CI ran 3.10/3.11/3.12 while wheels shipped cp310–cp314, so the two newest interpreters
were published untested — and four distros in the test matrix already ship 3.14. CI is now
3.12/3.13/3.14 and the wheels are cp312–cp314, so every published interpreter is tested.
requires-python is >=3.12, which means Ubuntu 22.04 LTS (3.10) and Debian 12 (3.11)
can no longer install glq. pip will say so at resolve time rather than handing over an
untested wheel. Every distro in the matrix still clears the floor — Ubuntu 24.04 → 3.12,
Debian 13 → 3.13, RHEL 9 family → 3.12, Fedora/Arch/Azure Linux → 3.14, Tumbleweed → 3.13.
Fixes
- The MoE kernel refused nothing. Shapes off its contract (
m % 32,k % 64) came
back finite, plausible and wrong;hidden=96is a legal MoE width that packs fine and
hit exactly that. It errors now, as the dense entry always has. - An in-place write crashed vLLM on the first token. Every forward runs inside
InferenceMode, and ATen's parallel workers do not inherit that state, so the scratch
scatter raised "Inplace update to inference tensor outside InferenceMode" and ended
EngineCore. Raw-pointer scatter instead — correct in either mode, on any thread. numpywas never a declared dependency, thoughglq.trellisimports it at module
scope and the vLLM plugin importsglq.trelliswhile loading weights. A bare
pip install glqtherefore could not serve the default codebook; both usual installs
masked it by pulling numpy in through vLLM or transformers.- The kernel loader told CPU-only users their CUDA build had "failed" and that trellis
vLLM serving "will not work" — the thing they are doing. That case now names the CPU
path; a machine that has a device and still failed to build keeps the full diagnosis.
Measured
8-vCPU Sapphire Rapids (AVX-512-FP16, 4 physical cores), vLLM 0.28.0+cpu, one request at
a time, 64 decoded tokens. These are single-stream numbers on one machine, not a general
claim.
| decode | footprint | |
|---|---|---|
gemma-4-26B-A4B-it-GLQ-trellis-3inst-4bpw, vLLM's default core binding |
3.0 tok/s | 18.6 GiB resident (13.9 GiB weights + 4 GiB KV), load 81 s, TTFT 0.4 s |
same, reserved core claimed (what glq-chat does) |
3.4 tok/s | |
| same, per-expert loop instead of the fused op | 2.8 tok/s | |
dense gemma-4-E4B-it-GLQ-trellis-3inst-4bpw, same box |
2.7 tok/s | 6.1 GiB |
Also measured and deliberately not adopted: --async-scheduling is neutral on CPU;
--kv-cache-dtype fp8 doubles KV capacity (18,533 → 37,067 tokens) at no throughput cost
but is a quality change with no quality measurement here; n-gram speculation is slower
(2.6 tok/s).
Validation
20 MoE kernel gates and 119 dense gates across all four ISA tiers on that box, including
bit-exactness against a pure-torch oracle. The vLLM-level tests assert the mechanism — the
fused entry is called exactly once for a layer inside the gate and not at all for one
outside it — because both paths agree numerically and output alone cannot say which ran.
Full local suite 1407 passed. install.sh pre-flight verified in pristine ubuntu:24.04,
debian:13, fedora:latest and opensuse/tumbleweed containers, plus a from-source --cpu
install that compiles the CPU extension on a clean image.
The distro matrix's GPU legs were not run — they need --gpus all and no GPU box was
available. Existing checkpoints are unaffected: no CUDA path, checkpoint format or
quantize path changed.