Skip to content

v20261001.1 — vLLM 0.29.0+rv.mxfp8.mtp.0a6739845a32.965e748f74fe.c6f1c3f3e582

Latest

Choose a tag to compare

@randomvariable randomvariable released this 01 Oct 00:03
· 1 commit to main since this release
607a961

Linux arm64 (SM12x family) and amd64 (sm_120) image index, vLLM 0.29.0+rv.mxfp8.mtp.0a6739845a32.965e748f74fe.c6f1c3f3e582.

docker pull ghcr.io/randomvariable/vllm-b12x-multi@sha256:43aaf6d1b5e0b96bff630f5b5211afccd1b4838c40a29864aaee3eb7dae6cbde
Field Value
Digest sha256:43aaf6d1b5e0b96bff630f5b5211afccd1b4838c40a29864aaee3eb7dae6cbde
Release tag ghcr.io/randomvariable/vllm-b12x-multi:v20261001.1
Publication tag ghcr.io/randomvariable/vllm-b12x-multi:vllmb12x-dev-rv-mxfp8-mtp-0a6739845a32-607a9619a9c3-20261001-n70
vLLM 0a6739845a3249b07a30ad9b1e720e5e3fb6236f
B12X 965e748f74fe9dac0a35c497a0947c5b108617c5
Builder 607a9619a9c31e7b75d776e0ee5d398b8067fe2e

Base vLLM image additions

  • Mooncake Transfer Engine CUDA 13 0.3.13.post1 provides the mooncake Python package and its service and benchmark commands for deployments that select vLLM's MooncakeStoreConnector. It does not enable a connector or start Mooncake services.
  • See Mooncake Transfer Engine for the installed surface and deployment boundary.

Included upstream changes

The image is built from pinned fork revisions rather than from upstream branches, so this table records which upstream changes the current lock carries. Regenerate it with scripts/vllmb12x-included-changes.py.

Component Change Included as
vLLM 0a6739845a32 Native MXFP8 MTP through the ModelOpt B12X backend (no upstream pull request) Merged into the pinned revision (4 commits)
vLLM 0a6739845a32 Partial port of vLLM PR 779: Qwen HC ownership and checkpoint coalescing (no upstream pull request) Merged into the pinned revision (9 commits)
vLLM 0a6739845a32 local-inference-lab/vllm#800 Bounded shared-memory broadcast waits Merged into the pinned revision (1 commits)
vLLM 0a6739845a32 Rewrite flash_attn.cute imports to vllm.vllm_flash_attn.cute (no upstream pull request) Applied at build time by third_party/vllm_flash_attn_cute_namespace.patch
B12X 965e748f74fe Native block-scaled MXFP8 W8A8 MoE execution (no upstream pull request) Merged into the pinned revision (3 commits)
B12X 965e748f74fe PLE checkpoint exports, prepared launcher closures and collective entry barriers (no upstream pull request) Merged into the pinned revision (11 commits)

Serving verification (2 GB10 groups x TP=2, idle, image above)

llm-inference-bench 0.7.5 against one group's rank-zero endpoint, 30 s cells, one sample per cell, temperature 1, max 8,192 output tokens. Qwen3.8-Flash-Next-NVFP4 qad-step5500-ple1000, MTP depth 3, --max-num-seqs 8, utilization 0.74.

  • Backend selection: B12X (NVFP4 target experts), B12X_MXFP8 (MXFP8 draft experts); nothing on Marlin. 72/72 B12X preparation units ready.
  • Prefill: 2,857 / 2,882 / 2,664 / 2,416 tok/s at 8K / 16K / 64K / 128K (TTFT 2.87 / 5.61 / 24.07 / 52.98 s). The 128K cell ranged 1,873-2,416 tok/s across three single samples.
  • Decode, aggregate tok/s / engine steps/s (accept length): C1 52.5 / 23.4 (2.24); C4 126.7 / 59.3 (2.14); C8 181.8 / 86.8 (2.09). Context 0 and 16K agree within ~3% on steps/s.
  • Not measured: gateway-path overhead, prefix-cache reuse, comparison against Marlin.