Repository navigation
Releases: lukolszewski/llama.cpp-multigpu
Release list
llama.cpp-multigpu 20261006.1 — Qwen3.8-Flash-Next multi-GPU performance build
llama.cpp-multigpu
Upstream llama.cpp plus 47 performance patches for Qwen3.8-Flash-Next (qwen4exp) on consumer multi-GPU systems with
constrained PCIe (measured on 6 × RTX 3090). Same engine, same GGUF models, same llama-server API; the whole patchset is
selected at runtime by environment variables and a few server flags, so one build carries all of it. This is the
configuration running in production on the benchmark machine since 2026-10-05.
Measured against upstream (same machine, same command line; upstream = the commit this patchset branches from)
| workload | 5k | 50k | 150k | 200k | 250k tokens |
|---|---|---|---|---|---|
| 1 slot prefill, upstream → multigpu (t/s) | 715 → 1249 | 568 → 2127 | 368 → 2152 | 315 → 2122 | 263 → 2111 |
| 1 slot decode (t/s) | 38.5 → 45.9 | 25.7 → 41.9 | 14.1 → 37.8 | 11.9 → 36.6 | 10.2 → 33.7 |
| 5 slots prefill, aggregate (t/s) | 630 → 2120 | 574 → 2301 | 363 → 2251 | 307 → 2180 | 265 → 2060 |
| 5 slots decode, per slot (t/s) | 18.2 → 42.3 | 8.1 → 38.3 | 3.6 → 32.4 | 2.8 → 29.0 | 2.3 → 27.3 |
Protocol, configuration and raw data: docs/multigpu/benchmarks.md, benches/multi-gpu/machine-01-7950x-6x3090/.
What the patches do (every switch and commit: docs/multigpu/patches.md)
- Long-context decode and quantized KV (patches 1–2). Fixes the CUDA decode slowdown of
qwen4expat depth and adds
a compact-gather attention path so aq8_0KV cache keeps the sparse-attention speedup (upstream falls back to a dense
path): 225k-context decode 22 → 44 t/s. - Server scheduling (3–4). Busy slots stay on contiguous sequence ids so several sessions decode in one ubatch; a
prefill admission policy bounds how many prompts prefill at once (--prefill-max-partial,--prefill-max-long). - Pipeline parallelism and device-built inputs (5–8). Ubatches overlap across GPUs for this hybrid model
(LLAMA_PIPELINE_PARALLEL=1,GGML_CUDA_GRAPHS_FORCE=1); the causal KQ mask and the QSA block bias are built on the
device instead of being copied over PCIe every ubatch: prefill at 82k 1260 → 2360 t/s, flat to 250k. - Per-sequence decode pipeline and decode groups (9–27).
LLAMA_DECODE_PIPELINE=1splits multi-user decode into
per-sequence ubatches on a stream-agnostic graph;LLAMA_SERVER_GROUPS=5keeps several groups' batches in flight so
users' tokens overlap across the GPUs instead of waiting for each other. 5 users at 58k: 17 → 34 t/s each. - Prefill while others decode (28–36). Prompts are prefilled in 256-token chunks beside the decoding slots without
draining the pipeline: two users decoding during a 200k prefill see ~280 ms between tokens instead of 1.3 s. - Asynchronous prompt cache (37–43). Slot save/restore of multi-GiB KV states runs on worker threads with pinned
staging and fences; context checkpoints no longer drain the pipeline; a 2.4 GiB restore takes ~1 s while the other
users keep decoding (worst hitch 0.4 s). - n-gram lookup speculation with the groups (44–47). Prompt-lookup decoding (no draft model) made to work with the
decode groups through recurrent-state snapshots in the memory, with a cap on drafting users and an acceptance gate.
Off by default;--spec-type ngram-map-k4v+LLAMA_SERVER_SPEC_MAX_USERS=2: solo code rewrite 44 → 111 t/s.
Also fixed on the way: an output-slot leak that crashed the server after client disconnects under load, an output-buffer
reallocation that corrupted parked results, and the upstream n-gram map keeping stale state across prompts.
Reference configuration (the numbers above were taken with it)
LLAMA_DECODE_PIPELINE=1 LLAMA_SERVER_GROUPS=5 LLAMA_PIPELINE_PARALLEL=1 GGML_CUDA_GRAPHS_FORCE=1 LLAMA_ATTN_ROT_DISABLE=1 \
llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -ngl 99 -fa on -np 5 -c 1310720 \
--cache-type-k q8_0 --cache-type-v q8_0 -b 2048 -ub 512 -ot 'per_layer_token_embd\.weight=CPU' \
--tensor-split 0.85,1,1,1,1,1 --prefill-max-partial 2
Scope and caveats
Targeted at one model family on one class of hardware; other models and NVLink or single-GPU machines may not benefit and
can regress. Only sm_86 is runtime-validated by the maintainer; the other architectures in the archives are compiled, not
run. Mixed prefill + decode is better than upstream but still a documented limitation. Upstream llama.cpp remains the
general-purpose project; this fork exists to be archived once upstream catches up.
Container images: forward-compatibility library removed (from multigpu-20261006.1)
NVIDIA's cuda:*-runtime base images ship cuda-compat-<ver>, a newer libcuda for datacenter GPUs on old
drivers. On a host whose driver is older than that library the container toolkit uses it, GeForce cards reject it
(ggml_cuda_init: failed to initialize CUDA: forward compatibility was attempted on non supported HW) and
llama.cpp silently runs on the CPU. Found on an 8 × RTX 4090 host with driver 570 using the multigpu-20261006
images; the package is purged from every image from this release on, so the host driver's libcuda is used
(any R525+ driver for the native targets). If you must run the multigpu-20261006 images on such a host:
-v /dev/null:/usr/local/cuda/compat/libcuda.so.1. The tarballs were never affected. Details:
docs/multigpu/builds.md.
This build
| release tag | multigpu-20261006.1 |
| multigpu commit | 79a12ee4767d (branch multigpu) |
| upstream llama.cpp base | df03399b8858 (2026-09-10) |
| commits ahead of upstream | 94 (47 code patches indexed in docs/multigpu/patches.md, the rest docs/CI; full list below, also attached as -patches.tar.gz) |
toolchain cuda-12.9 |
CUDA 12.9.1 (release 12.9), cc (Ubuntu 13.3.0-6ubuntu2~24.04) 13.3.0, cmake version 3.28.3, Ubuntu 24.04.2 LTS |
toolchain cuda-13.4 |
CUDA 13.4.1 (release 13.4), cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0, cmake version 3.28.3, Ubuntu 24.04.5 LTS |
Downloads
Two Linux x86-64 CUDA flavours. Pick by GPU generation and installed driver; both carry the same code.
| asset | CUDA | GPUs with native code (SASS) | PTX for other GPUs (JIT by the driver) | minimum driver | validation |
|---|---|---|---|---|---|
llama.cpp-multigpu-20261006.1-bin-ubuntu-cuda-12.9-x64.tar.gz |
12.9.1 | sm_70, sm_86, sm_89, sm_120a, sm_121a | sm_50, sm_61, sm_70, sm_75, sm_80, sm_90 | R525+ (native targets); R570+ for Blackwell; R575+ for the PTX-only GPUs (CUDA 12.9 PTX JIT) | Tier A passed; Tier B (sm_86, 6 x RTX 3090) by the maintainer |
llama.cpp-multigpu-20261006.1-cudart-ubuntu-cuda-12.9-x64.tar.gz |
12.9.1 | runtime libraries only | - | R525+ (native targets); R570+ for Blackwell; R575+ for the PTX-only GPUs (CUDA 12.9 PTX JIT) | - |
llama.cpp-multigpu-20261006.1-bin-ubuntu-cuda-13.4-x64.tar.gz |
13.4.1 | sm_86, sm_89, sm_120a, sm_121a | sm_80, sm_90 | R580+ (CUDA 13); a driver as new as CUDA 13.4 for the PTX-only GPUs | Tier A passed; not runtime-validated by us |
llama.cpp-multigpu-20261006.1-cudart-ubuntu-cuda-13.4-x64.tar.gz |
13.4.1 | runtime libraries only | - | R580+ (CUDA 13); a driver as new as CUDA 13.4 for the PTX-only GPUs | - |
llama.cpp-multigpu-20261006.1-patches.tar.gz |
- | git format-patch series | - | - | - |
-
cuda-12.9: the wide build. Native code for Volta (V100,sm_70), Ampere (sm_86), Ada (sm_89) and Blackwell
(sm_120a/sm_121a); PTX for Maxwell, Pascal, Turing, A100 (sm_80) and Hopper (sm_90), which the driver compiles
at first load. Needs an R525+ driver for the native targets, R570+ for Blackwell cards, and a driver that understands
CUDA 12.9 PTX (R575+) for the PTX-only GPUs. -
cuda-13.4: Ampere and newer only (CUDA 13 dropped Maxwell/Pascal/Volta compilation). Nativesm_86/sm_89/
sm_120a/sm_121a, PTX forsm_80/sm_90. Needs an R580+ driver. -
cudart-*: the CUDA runtime + cuBLAS libraries matching each flavour, for hosts without a CUDA toolkit. Unpack next to
the binaries (or setLD_LIBRARY_PATH). The NVIDIA driver (libcuda.so.1) always comes from the host. -
patches.tar.gz: the patch series asgit format-patchoutput, for building upstream + selected patches. -
Container images (public, anonymous pull; built from these very archives;
llama-serveris the entrypoint, run
with--gpus all) — package page: https://github.com/lukolszewski/llama.cpp-multigpu/pkgs/container/llama.cpp-multigpuimage contents docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:server-cuda12.9-20261006.1cuda-12.9 archive on nvidia/cuda:12.9.1-runtime-ubuntu24.04; V100 → RTX 5090; moving tags:server-cuda12.9,:latestdocker pull ghcr.io/lukolszewski/llama.cpp-multigpu:server-cuda13.4-20261006.1cuda-13.4 archive on nvidia/cuda:13.4.1-runtime-ubuntu24.04; Ampere and newer; moving tag:server-cuda13.4; the base image requires a driver reporting CUDA ≥ 13.4 (-e NVIDIA_DISABLE_REQUIRE=1on 13.x drivers)docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:bench-cuda12.9-20261006.1the cuda-12.9 server image plus the README benchmark grid runner for rented machines ( scripts/multigpu/bench/vast/README.md); moving tag:bench-cuda12.9Upstream baseline used for the comparisons: https://github.com/lukolszewski/llama.cpp-multigpu/releases/tag/baseline-df03399
(unmodified upstreamdf03399b8, same cuda-12.9 recipe, tarball only).
Validation: Tier A = the offline packaging gate passed in CI (archive unpacks, binaries run, the declared SASS/PTX
targets are really embedded, the downstream flags are wired, metadata matches the commit). Tier B ...
Upstream baseline df03399b8 (b10902), cuda-12.9 recipe
Unmodified upstream llama.cpp commit df03399b8 (build b10902), compiled with the cuda-12.9 recipe of this repository's release workflow (CUDA 12.9.1, SASS sm_70/86/89/120a/121a, PTX sm_50/61/70/75/80/90), no downstream patches.
It is the upstream side of the benchmark grids in benches/multi-gpu/ and is fetched by the bench image (scripts/multigpu/bench/vast/). Not a release of llama.cpp-multigpu; for those see the multigpu-* releases. Needs the CUDA 12.9 runtime libraries (cudart-… archive of any multigpu-* release, or NVIDIA's nvidia/cuda:12.9.1-runtime image) and an R525+ driver (R575+ for PTX-only GPUs).