Skip to content

Releases: randomvariable/vllm-multiarch-oci

v20261001.1 — vLLM 0.29.0+rv.mxfp8.mtp.0a6739845a32.965e748f74fe.c6f1c3f3e582

Choose a tag to compare

@randomvariable randomvariable released this 01 Oct 00:03
607a961

Linux arm64 (SM12x family) and amd64 (sm_120) image index, vLLM 0.29.0+rv.mxfp8.mtp.0a6739845a32.965e748f74fe.c6f1c3f3e582.

docker pull ghcr.io/randomvariable/vllm-b12x-multi@sha256:43aaf6d1b5e0b96bff630f5b5211afccd1b4838c40a29864aaee3eb7dae6cbde
Field Value
Digest sha256:43aaf6d1b5e0b96bff630f5b5211afccd1b4838c40a29864aaee3eb7dae6cbde
Release tag ghcr.io/randomvariable/vllm-b12x-multi:v20261001.1
Publication tag ghcr.io/randomvariable/vllm-b12x-multi:vllmb12x-dev-rv-mxfp8-mtp-0a6739845a32-607a9619a9c3-20261001-n70
vLLM 0a6739845a3249b07a30ad9b1e720e5e3fb6236f
B12X 965e748f74fe9dac0a35c497a0947c5b108617c5
Builder 607a9619a9c31e7b75d776e0ee5d398b8067fe2e

Base vLLM image additions

  • Mooncake Transfer Engine CUDA 13 0.3.13.post1 provides the mooncake Python package and its service and benchmark commands for deployments that select vLLM's MooncakeStoreConnector. It does not enable a connector or start Mooncake services.
  • See Mooncake Transfer Engine for the installed surface and deployment boundary.

Included upstream changes

The image is built from pinned fork revisions rather than from upstream branches, so this table records which upstream changes the current lock carries. Regenerate it with scripts/vllmb12x-included-changes.py.

Component Change Included as
vLLM 0a6739845a32 Native MXFP8 MTP through the ModelOpt B12X backend (no upstream pull request) Merged into the pinned revision (4 commits)
vLLM 0a6739845a32 Partial port of vLLM PR 779: Qwen HC ownership and checkpoint coalescing (no upstream pull request) Merged into the pinned revision (9 commits)
vLLM 0a6739845a32 local-inference-lab/vllm#800 Bounded shared-memory broadcast waits Merged into the pinned revision (1 commits)
vLLM 0a6739845a32 Rewrite flash_attn.cute imports to vllm.vllm_flash_attn.cute (no upstream pull request) Applied at build time by third_party/vllm_flash_attn_cute_namespace.patch
B12X 965e748f74fe Native block-scaled MXFP8 W8A8 MoE execution (no upstream pull request) Merged into the pinned revision (3 commits)
B12X 965e748f74fe PLE checkpoint exports, prepared launcher closures and collective entry barriers (no upstream pull request) Merged into the pinned revision (11 commits)

Serving verification (2 GB10 groups x TP=2, idle, image above)

llm-inference-bench 0.7.5 against one group's rank-zero endpoint, 30 s cells, one sample per cell, temperature 1, max 8,192 output tokens. Qwen3.8-Flash-Next-NVFP4 qad-step5500-ple1000, MTP depth 3, --max-num-seqs 8, utilization 0.74.

  • Backend selection: B12X (NVFP4 target experts), B12X_MXFP8 (MXFP8 draft experts); nothing on Marlin. 72/72 B12X preparation units ready.
  • Prefill: 2,857 / 2,882 / 2,664 / 2,416 tok/s at 8K / 16K / 64K / 128K (TTFT 2.87 / 5.61 / 24.07 / 52.98 s). The 128K cell ranged 1,873-2,416 tok/s across three single samples.
  • Decode, aggregate tok/s / engine steps/s (accept length): C1 52.5 / 23.4 (2.24); C4 126.7 / 59.3 (2.14); C8 181.8 / 86.8 (2.09). Context 0 and 16K agree within ~3% on steps/s.
  • Not measured: gateway-path overhead, prefix-cache reuse, comparison against Marlin.

v20260917.1 — vLLM 0.1.dev1+vllmb12x.gbd22e0f25043

Choose a tag to compare

@randomvariable randomvariable released this 17 Sep 11:37
13527e9

Linux ARM64 image for NVIDIA DGX Spark (GB10, sm_121a), vLLM 0.1.dev1+vllmb12x.gbd22e0f25043.

docker pull ghcr.io/randomvariable/vllm-b12x-multi@sha256:cdf2602bb2e42da5b8f8cae8ae4bbe0842388cd6d8bf6ca2e6e3c27d7e27ad1e
Field Value
Digest sha256:cdf2602bb2e42da5b8f8cae8ae4bbe0842388cd6d8bf6ca2e6e3c27d7e27ad1e
Release tag ghcr.io/randomvariable/vllm-b12x-multi:v20260917.1
Publication tag ghcr.io/randomvariable/vllm-b12x-multi:vllmb12x-dev-rv-jovian-judgement-profile-base-bd22e0f25043-13527e984717-20260917-n59
vLLM bd22e0f250439659d9fe30903b7aa124c321a932
B12X 4401ce4fb98265c3b1e1579ae03960ff41d84d50
Builder 13527e9847178f123098b7967c27333f4c82773b

Included upstream changes

The image is built from pinned fork revisions rather than from upstream branches, so this table records which upstream changes the current lock carries. Regenerate it with scripts/vllmb12x-included-changes.py.

Component Change Included as
vLLM bd22e0f25043 local-inference-lab/vllm#777 fix(qwen): propagate MTP positional overrides Merged into the pinned revision (3 commits)
vLLM bd22e0f25043 local-inference-lab/vllm#779 perf(qwen): shard TP4 HC prefill and coalesce recurrent checkpoints Merged into the pinned revision (9 commits)
vLLM bd22e0f25043 vllm-project/vllm#52917 Adaptive spin grace and bounded architectural waits for shm_broadcast Applied at build time by third_party/vllm_shm_broadcast_spin_grace.patch
vLLM bd22e0f25043 Rewrite flash_attn.cute imports to vllm.vllm_flash_attn.cute (no upstream pull request) Applied at build time by third_party/vllm_flash_attn_cute_namespace.patch
B12X 4401ce4fb982 local-inference-lab/b12x#384 fix(preparation): retain prepared launchers and coordinate collectives Merged into the pinned revision (10 commits)
B12X 4401ce4fb982 local-inference-lab/b12x#386 feat(ple): export prepared internal prefill checkpoints Merged into the pinned revision (2 commits)
B12X 4401ce4fb982 local-inference-lab/b12x#387 perf(qsa): reuse representative keys across paired queries Merged into the pinned revision (4 commits)