Skip to content

v20260917.1 — vLLM 0.1.dev1+vllmb12x.gbd22e0f25043

Choose a tag to compare

@randomvariable randomvariable released this 17 Sep 11:37
· 18 commits to main since this release
13527e9

Linux ARM64 image for NVIDIA DGX Spark (GB10, sm_121a), vLLM 0.1.dev1+vllmb12x.gbd22e0f25043.

docker pull ghcr.io/randomvariable/vllm-b12x-multi@sha256:cdf2602bb2e42da5b8f8cae8ae4bbe0842388cd6d8bf6ca2e6e3c27d7e27ad1e
Field Value
Digest sha256:cdf2602bb2e42da5b8f8cae8ae4bbe0842388cd6d8bf6ca2e6e3c27d7e27ad1e
Release tag ghcr.io/randomvariable/vllm-b12x-multi:v20260917.1
Publication tag ghcr.io/randomvariable/vllm-b12x-multi:vllmb12x-dev-rv-jovian-judgement-profile-base-bd22e0f25043-13527e984717-20260917-n59
vLLM bd22e0f250439659d9fe30903b7aa124c321a932
B12X 4401ce4fb98265c3b1e1579ae03960ff41d84d50
Builder 13527e9847178f123098b7967c27333f4c82773b

Included upstream changes

The image is built from pinned fork revisions rather than from upstream branches, so this table records which upstream changes the current lock carries. Regenerate it with scripts/vllmb12x-included-changes.py.

Component Change Included as
vLLM bd22e0f25043 local-inference-lab/vllm#777 fix(qwen): propagate MTP positional overrides Merged into the pinned revision (3 commits)
vLLM bd22e0f25043 local-inference-lab/vllm#779 perf(qwen): shard TP4 HC prefill and coalesce recurrent checkpoints Merged into the pinned revision (9 commits)
vLLM bd22e0f25043 vllm-project/vllm#52917 Adaptive spin grace and bounded architectural waits for shm_broadcast Applied at build time by third_party/vllm_shm_broadcast_spin_grace.patch
vLLM bd22e0f25043 Rewrite flash_attn.cute imports to vllm.vllm_flash_attn.cute (no upstream pull request) Applied at build time by third_party/vllm_flash_attn_cute_namespace.patch
B12X 4401ce4fb982 local-inference-lab/b12x#384 fix(preparation): retain prepared launchers and coordinate collectives Merged into the pinned revision (10 commits)
B12X 4401ce4fb982 local-inference-lab/b12x#386 feat(ple): export prepared internal prefill checkpoints Merged into the pinned revision (2 commits)
B12X 4401ce4fb982 local-inference-lab/b12x#387 perf(qsa): reuse representative keys across paired queries Merged into the pinned revision (4 commits)