Releases: randomvariable/vllm-multiarch-oci
Release list
v20261001.1 — vLLM 0.29.0+rv.mxfp8.mtp.0a6739845a32.965e748f74fe.c6f1c3f3e582
Linux arm64 (SM12x family) and amd64 (sm_120) image index, vLLM 0.29.0+rv.mxfp8.mtp.0a6739845a32.965e748f74fe.c6f1c3f3e582.
docker pull ghcr.io/randomvariable/vllm-b12x-multi@sha256:43aaf6d1b5e0b96bff630f5b5211afccd1b4838c40a29864aaee3eb7dae6cbde| Field | Value |
|---|---|
| Digest | sha256:43aaf6d1b5e0b96bff630f5b5211afccd1b4838c40a29864aaee3eb7dae6cbde |
| Release tag | ghcr.io/randomvariable/vllm-b12x-multi:v20261001.1 |
| Publication tag | ghcr.io/randomvariable/vllm-b12x-multi:vllmb12x-dev-rv-mxfp8-mtp-0a6739845a32-607a9619a9c3-20261001-n70 |
| vLLM | 0a6739845a3249b07a30ad9b1e720e5e3fb6236f |
| B12X | 965e748f74fe9dac0a35c497a0947c5b108617c5 |
| Builder | 607a9619a9c31e7b75d776e0ee5d398b8067fe2e |
Base vLLM image additions
- Mooncake Transfer Engine CUDA 13
0.3.13.post1provides themooncakePython package and its service and benchmark commands for deployments that select vLLM'sMooncakeStoreConnector. It does not enable a connector or start Mooncake services. - See Mooncake Transfer Engine for the installed surface and deployment boundary.
Included upstream changes
The image is built from pinned fork revisions rather than from upstream branches, so this table records which upstream changes the current lock carries. Regenerate it with scripts/vllmb12x-included-changes.py.
| Component | Change | Included as |
|---|---|---|
vLLM 0a6739845a32 |
Native MXFP8 MTP through the ModelOpt B12X backend (no upstream pull request) | Merged into the pinned revision (4 commits) |
vLLM 0a6739845a32 |
Partial port of vLLM PR 779: Qwen HC ownership and checkpoint coalescing (no upstream pull request) | Merged into the pinned revision (9 commits) |
vLLM 0a6739845a32 |
local-inference-lab/vllm#800 Bounded shared-memory broadcast waits | Merged into the pinned revision (1 commits) |
vLLM 0a6739845a32 |
Rewrite flash_attn.cute imports to vllm.vllm_flash_attn.cute (no upstream pull request) | Applied at build time by third_party/vllm_flash_attn_cute_namespace.patch |
B12X 965e748f74fe |
Native block-scaled MXFP8 W8A8 MoE execution (no upstream pull request) | Merged into the pinned revision (3 commits) |
B12X 965e748f74fe |
PLE checkpoint exports, prepared launcher closures and collective entry barriers (no upstream pull request) | Merged into the pinned revision (11 commits) |
Serving verification (2 GB10 groups x TP=2, idle, image above)
llm-inference-bench 0.7.5 against one group's rank-zero endpoint, 30 s cells, one sample per cell, temperature 1, max 8,192 output tokens. Qwen3.8-Flash-Next-NVFP4 qad-step5500-ple1000, MTP depth 3, --max-num-seqs 8, utilization 0.74.
- Backend selection:
B12X(NVFP4 target experts),B12X_MXFP8(MXFP8 draft experts); nothing on Marlin. 72/72 B12X preparation units ready. - Prefill: 2,857 / 2,882 / 2,664 / 2,416 tok/s at 8K / 16K / 64K / 128K (TTFT 2.87 / 5.61 / 24.07 / 52.98 s). The 128K cell ranged 1,873-2,416 tok/s across three single samples.
- Decode, aggregate tok/s / engine steps/s (accept length): C1 52.5 / 23.4 (2.24); C4 126.7 / 59.3 (2.14); C8 181.8 / 86.8 (2.09). Context 0 and 16K agree within ~3% on steps/s.
- Not measured: gateway-path overhead, prefix-cache reuse, comparison against Marlin.
v20260917.1 — vLLM 0.1.dev1+vllmb12x.gbd22e0f25043
Linux ARM64 image for NVIDIA DGX Spark (GB10, sm_121a), vLLM 0.1.dev1+vllmb12x.gbd22e0f25043.
docker pull ghcr.io/randomvariable/vllm-b12x-multi@sha256:cdf2602bb2e42da5b8f8cae8ae4bbe0842388cd6d8bf6ca2e6e3c27d7e27ad1e| Field | Value |
|---|---|
| Digest | sha256:cdf2602bb2e42da5b8f8cae8ae4bbe0842388cd6d8bf6ca2e6e3c27d7e27ad1e |
| Release tag | ghcr.io/randomvariable/vllm-b12x-multi:v20260917.1 |
| Publication tag | ghcr.io/randomvariable/vllm-b12x-multi:vllmb12x-dev-rv-jovian-judgement-profile-base-bd22e0f25043-13527e984717-20260917-n59 |
| vLLM | bd22e0f250439659d9fe30903b7aa124c321a932 |
| B12X | 4401ce4fb98265c3b1e1579ae03960ff41d84d50 |
| Builder | 13527e9847178f123098b7967c27333f4c82773b |
Included upstream changes
The image is built from pinned fork revisions rather than from upstream branches, so this table records which upstream changes the current lock carries. Regenerate it with scripts/vllmb12x-included-changes.py.
| Component | Change | Included as |
|---|---|---|
vLLM bd22e0f25043 |
local-inference-lab/vllm#777 fix(qwen): propagate MTP positional overrides | Merged into the pinned revision (3 commits) |
vLLM bd22e0f25043 |
local-inference-lab/vllm#779 perf(qwen): shard TP4 HC prefill and coalesce recurrent checkpoints | Merged into the pinned revision (9 commits) |
vLLM bd22e0f25043 |
vllm-project/vllm#52917 Adaptive spin grace and bounded architectural waits for shm_broadcast | Applied at build time by third_party/vllm_shm_broadcast_spin_grace.patch |
vLLM bd22e0f25043 |
Rewrite flash_attn.cute imports to vllm.vllm_flash_attn.cute (no upstream pull request) | Applied at build time by third_party/vllm_flash_attn_cute_namespace.patch |
B12X 4401ce4fb982 |
local-inference-lab/b12x#384 fix(preparation): retain prepared launchers and coordinate collectives | Merged into the pinned revision (10 commits) |
B12X 4401ce4fb982 |
local-inference-lab/b12x#386 feat(ple): export prepared internal prefill checkpoints | Merged into the pinned revision (2 commits) |
B12X 4401ce4fb982 |
local-inference-lab/b12x#387 perf(qsa): reuse representative keys across paired queries | Merged into the pinned revision (4 commits) |