Linux arm64 (SM12x family) and amd64 (sm_120) image index, vLLM 0.29.0+rv.mxfp8.mtp.0a6739845a32.965e748f74fe.c6f1c3f3e582.
docker pull ghcr.io/randomvariable/vllm-b12x-multi@sha256:43aaf6d1b5e0b96bff630f5b5211afccd1b4838c40a29864aaee3eb7dae6cbde| Field | Value |
|---|---|
| Digest | sha256:43aaf6d1b5e0b96bff630f5b5211afccd1b4838c40a29864aaee3eb7dae6cbde |
| Release tag | ghcr.io/randomvariable/vllm-b12x-multi:v20261001.1 |
| Publication tag | ghcr.io/randomvariable/vllm-b12x-multi:vllmb12x-dev-rv-mxfp8-mtp-0a6739845a32-607a9619a9c3-20261001-n70 |
| vLLM | 0a6739845a3249b07a30ad9b1e720e5e3fb6236f |
| B12X | 965e748f74fe9dac0a35c497a0947c5b108617c5 |
| Builder | 607a9619a9c31e7b75d776e0ee5d398b8067fe2e |
Base vLLM image additions
- Mooncake Transfer Engine CUDA 13
0.3.13.post1provides themooncakePython package and its service and benchmark commands for deployments that select vLLM'sMooncakeStoreConnector. It does not enable a connector or start Mooncake services. - See Mooncake Transfer Engine for the installed surface and deployment boundary.
Included upstream changes
The image is built from pinned fork revisions rather than from upstream branches, so this table records which upstream changes the current lock carries. Regenerate it with scripts/vllmb12x-included-changes.py.
| Component | Change | Included as |
|---|---|---|
vLLM 0a6739845a32 |
Native MXFP8 MTP through the ModelOpt B12X backend (no upstream pull request) | Merged into the pinned revision (4 commits) |
vLLM 0a6739845a32 |
Partial port of vLLM PR 779: Qwen HC ownership and checkpoint coalescing (no upstream pull request) | Merged into the pinned revision (9 commits) |
vLLM 0a6739845a32 |
local-inference-lab/vllm#800 Bounded shared-memory broadcast waits | Merged into the pinned revision (1 commits) |
vLLM 0a6739845a32 |
Rewrite flash_attn.cute imports to vllm.vllm_flash_attn.cute (no upstream pull request) | Applied at build time by third_party/vllm_flash_attn_cute_namespace.patch |
B12X 965e748f74fe |
Native block-scaled MXFP8 W8A8 MoE execution (no upstream pull request) | Merged into the pinned revision (3 commits) |
B12X 965e748f74fe |
PLE checkpoint exports, prepared launcher closures and collective entry barriers (no upstream pull request) | Merged into the pinned revision (11 commits) |
Serving verification (2 GB10 groups x TP=2, idle, image above)
llm-inference-bench 0.7.5 against one group's rank-zero endpoint, 30 s cells, one sample per cell, temperature 1, max 8,192 output tokens. Qwen3.8-Flash-Next-NVFP4 qad-step5500-ple1000, MTP depth 3, --max-num-seqs 8, utilization 0.74.
- Backend selection:
B12X(NVFP4 target experts),B12X_MXFP8(MXFP8 draft experts); nothing on Marlin. 72/72 B12X preparation units ready. - Prefill: 2,857 / 2,882 / 2,664 / 2,416 tok/s at 8K / 16K / 64K / 128K (TTFT 2.87 / 5.61 / 24.07 / 52.98 s). The 128K cell ranged 1,873-2,416 tok/s across three single samples.
- Decode, aggregate tok/s / engine steps/s (accept length): C1 52.5 / 23.4 (2.24); C4 126.7 / 59.3 (2.14); C8 181.8 / 86.8 (2.09). Context 0 and 16K agree within ~3% on steps/s.
- Not measured: gateway-path overhead, prefix-cache reuse, comparison against Marlin.