Skip to content

v0.8.17 — trellis MoE serves on CPU

Choose a tag to compare

@cnygaard cnygaard released this 05 Sep 10:57
· 34 commits to main since this release
bd90d38

A trellis MoE checkpoint now serves on a machine with no GPU. Until 0.8.16 the vLLM
CPU backend refused every GLQ MoE, which put the most CPU-suited architecture we publish
out of reach: an MoE reads only its top-k experts per token, so it moves far less memory
than a dense model of the same size.

What changed

  • CPU MoE serving. The refusal is per codebook now — trellis is allowed, e8p/shell
    still refuse at load with the reason and the fix, because their expert decode is
    CUDA-only.
  • The fused CPU expert kernel runs under vLLM. GLQFusedMoEMethod.apply() takes a CPU
    branch: one extension call per MoE block instead of a Python loop of per-expert calls.
    The CUDA path is untouched. Anything outside the op's reach (stage-2 RVQ, padded shapes,
    ungated activations, an older wheel) falls back to the loop — correct on any shape — and
    says why, once, rather than dropping onto the slower path silently.
  • The installer recommends it. A trellis MoE that fits the RAM budget is now the CPU
    recommendation; e8p/shell and unknown formats are still excluded, and the budget is
    unchanged.
  • glq-chat claims the core vLLM holds back (VLLM_CPU_NUM_OF_RESERVED_CPU=0),
    worth ~11% on a 4-core box. Set it yourself to keep a core reserved.

Breaking: Python 3.12 is now the floor

CI ran 3.10/3.11/3.12 while wheels shipped cp310–cp314, so the two newest interpreters
were published untested — and four distros in the test matrix already ship 3.14. CI is now
3.12/3.13/3.14 and the wheels are cp312–cp314, so every published interpreter is tested.

requires-python is >=3.12, which means Ubuntu 22.04 LTS (3.10) and Debian 12 (3.11)
can no longer install glq
. pip will say so at resolve time rather than handing over an
untested wheel. Every distro in the matrix still clears the floor — Ubuntu 24.04 → 3.12,
Debian 13 → 3.13, RHEL 9 family → 3.12, Fedora/Arch/Azure Linux → 3.14, Tumbleweed → 3.13.

Fixes

  • The MoE kernel refused nothing. Shapes off its contract (m % 32, k % 64) came
    back finite, plausible and wrong; hidden=96 is a legal MoE width that packs fine and
    hit exactly that. It errors now, as the dense entry always has.
  • An in-place write crashed vLLM on the first token. Every forward runs inside
    InferenceMode, and ATen's parallel workers do not inherit that state, so the scratch
    scatter raised "Inplace update to inference tensor outside InferenceMode" and ended
    EngineCore. Raw-pointer scatter instead — correct in either mode, on any thread.
  • numpy was never a declared dependency, though glq.trellis imports it at module
    scope and the vLLM plugin imports glq.trellis while loading weights. A bare
    pip install glq therefore could not serve the default codebook; both usual installs
    masked it by pulling numpy in through vLLM or transformers.
  • The kernel loader told CPU-only users their CUDA build had "failed" and that trellis
    vLLM serving "will not work" — the thing they are doing. That case now names the CPU
    path; a machine that has a device and still failed to build keeps the full diagnosis.

Measured

8-vCPU Sapphire Rapids (AVX-512-FP16, 4 physical cores), vLLM 0.28.0+cpu, one request at
a time, 64 decoded tokens. These are single-stream numbers on one machine, not a general
claim.

decode footprint
gemma-4-26B-A4B-it-GLQ-trellis-3inst-4bpw, vLLM's default core binding 3.0 tok/s 18.6 GiB resident (13.9 GiB weights + 4 GiB KV), load 81 s, TTFT 0.4 s
same, reserved core claimed (what glq-chat does) 3.4 tok/s
same, per-expert loop instead of the fused op 2.8 tok/s
dense gemma-4-E4B-it-GLQ-trellis-3inst-4bpw, same box 2.7 tok/s 6.1 GiB

Also measured and deliberately not adopted: --async-scheduling is neutral on CPU;
--kv-cache-dtype fp8 doubles KV capacity (18,533 → 37,067 tokens) at no throughput cost
but is a quality change with no quality measurement here; n-gram speculation is slower
(2.6 tok/s).

Validation

20 MoE kernel gates and 119 dense gates across all four ISA tiers on that box, including
bit-exactness against a pure-torch oracle. The vLLM-level tests assert the mechanism — the
fused entry is called exactly once for a layer inside the gate and not at all for one
outside it — because both paths agree numerically and output alone cannot say which ran.
Full local suite 1407 passed. install.sh pre-flight verified in pristine ubuntu:24.04,
debian:13, fedora:latest and opensuse/tumbleweed containers, plus a from-source --cpu
install that compiles the CPU extension on a clean image.

The distro matrix's GPU legs were not run — they need --gpus all and no GPU box was
available. Existing checkpoints are unaffected: no CUDA path, checkpoint format or
quantize path changed.