Skip to content

Releases: lukolszewski/llama.cpp-multigpu

llama.cpp-multigpu 20261006.1 — Qwen3.8-Flash-Next multi-GPU performance build

Choose a tag to compare

@github-actions github-actions released this 06 Oct 19:23
79a12ee

llama.cpp-multigpu

Upstream llama.cpp plus 47 performance patches for Qwen3.8-Flash-Next (qwen4exp) on consumer multi-GPU systems with
constrained PCIe (measured on 6 × RTX 3090). Same engine, same GGUF models, same llama-server API; the whole patchset is
selected at runtime by environment variables and a few server flags, so one build carries all of it. This is the
configuration running in production on the benchmark machine since 2026-10-05.

Measured against upstream (same machine, same command line; upstream = the commit this patchset branches from)

workload 5k 50k 150k 200k 250k tokens
1 slot prefill, upstream → multigpu (t/s) 715 → 1249 568 → 2127 368 → 2152 315 → 2122 263 → 2111
1 slot decode (t/s) 38.5 → 45.9 25.7 → 41.9 14.1 → 37.8 11.9 → 36.6 10.2 → 33.7
5 slots prefill, aggregate (t/s) 630 → 2120 574 → 2301 363 → 2251 307 → 2180 265 → 2060
5 slots decode, per slot (t/s) 18.2 → 42.3 8.1 → 38.3 3.6 → 32.4 2.8 → 29.0 2.3 → 27.3

Protocol, configuration and raw data: docs/multigpu/benchmarks.md, benches/multi-gpu/machine-01-7950x-6x3090/.

What the patches do (every switch and commit: docs/multigpu/patches.md)

  1. Long-context decode and quantized KV (patches 1–2). Fixes the CUDA decode slowdown of qwen4exp at depth and adds
    a compact-gather attention path so a q8_0 KV cache keeps the sparse-attention speedup (upstream falls back to a dense
    path): 225k-context decode 22 → 44 t/s.
  2. Server scheduling (3–4). Busy slots stay on contiguous sequence ids so several sessions decode in one ubatch; a
    prefill admission policy bounds how many prompts prefill at once (--prefill-max-partial, --prefill-max-long).
  3. Pipeline parallelism and device-built inputs (5–8). Ubatches overlap across GPUs for this hybrid model
    (LLAMA_PIPELINE_PARALLEL=1, GGML_CUDA_GRAPHS_FORCE=1); the causal KQ mask and the QSA block bias are built on the
    device instead of being copied over PCIe every ubatch: prefill at 82k 1260 → 2360 t/s, flat to 250k.
  4. Per-sequence decode pipeline and decode groups (9–27). LLAMA_DECODE_PIPELINE=1 splits multi-user decode into
    per-sequence ubatches on a stream-agnostic graph; LLAMA_SERVER_GROUPS=5 keeps several groups' batches in flight so
    users' tokens overlap across the GPUs instead of waiting for each other. 5 users at 58k: 17 → 34 t/s each.
  5. Prefill while others decode (28–36). Prompts are prefilled in 256-token chunks beside the decoding slots without
    draining the pipeline: two users decoding during a 200k prefill see ~280 ms between tokens instead of 1.3 s.
  6. Asynchronous prompt cache (37–43). Slot save/restore of multi-GiB KV states runs on worker threads with pinned
    staging and fences; context checkpoints no longer drain the pipeline; a 2.4 GiB restore takes ~1 s while the other
    users keep decoding (worst hitch 0.4 s).
  7. n-gram lookup speculation with the groups (44–47). Prompt-lookup decoding (no draft model) made to work with the
    decode groups through recurrent-state snapshots in the memory, with a cap on drafting users and an acceptance gate.
    Off by default; --spec-type ngram-map-k4v + LLAMA_SERVER_SPEC_MAX_USERS=2: solo code rewrite 44 → 111 t/s.

Also fixed on the way: an output-slot leak that crashed the server after client disconnects under load, an output-buffer
reallocation that corrupted parked results, and the upstream n-gram map keeping stale state across prompts.

Reference configuration (the numbers above were taken with it)

LLAMA_DECODE_PIPELINE=1 LLAMA_SERVER_GROUPS=5 LLAMA_PIPELINE_PARALLEL=1 GGML_CUDA_GRAPHS_FORCE=1 LLAMA_ATTN_ROT_DISABLE=1 \
llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -ngl 99 -fa on -np 5 -c 1310720 \
  --cache-type-k q8_0 --cache-type-v q8_0 -b 2048 -ub 512 -ot 'per_layer_token_embd\.weight=CPU' \
  --tensor-split 0.85,1,1,1,1,1 --prefill-max-partial 2

Scope and caveats

Targeted at one model family on one class of hardware; other models and NVLink or single-GPU machines may not benefit and
can regress. Only sm_86 is runtime-validated by the maintainer; the other architectures in the archives are compiled, not
run. Mixed prefill + decode is better than upstream but still a documented limitation. Upstream llama.cpp remains the
general-purpose project; this fork exists to be archived once upstream catches up.

Container images: forward-compatibility library removed (from multigpu-20261006.1)

NVIDIA's cuda:*-runtime base images ship cuda-compat-<ver>, a newer libcuda for datacenter GPUs on old
drivers. On a host whose driver is older than that library the container toolkit uses it, GeForce cards reject it
(ggml_cuda_init: failed to initialize CUDA: forward compatibility was attempted on non supported HW) and
llama.cpp silently runs on the CPU. Found on an 8 × RTX 4090 host with driver 570 using the multigpu-20261006
images; the package is purged from every image from this release on, so the host driver's libcuda is used
(any R525+ driver for the native targets). If you must run the multigpu-20261006 images on such a host:
-v /dev/null:/usr/local/cuda/compat/libcuda.so.1. The tarballs were never affected. Details:
docs/multigpu/builds.md.

This build

release tag multigpu-20261006.1
multigpu commit 79a12ee4767d (branch multigpu)
upstream llama.cpp base df03399b8858 (2026-09-10)
commits ahead of upstream 94 (47 code patches indexed in docs/multigpu/patches.md, the rest docs/CI; full list below, also attached as -patches.tar.gz)
toolchain cuda-12.9 CUDA 12.9.1 (release 12.9), cc (Ubuntu 13.3.0-6ubuntu2~24.04) 13.3.0, cmake version 3.28.3, Ubuntu 24.04.2 LTS
toolchain cuda-13.4 CUDA 13.4.1 (release 13.4), cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0, cmake version 3.28.3, Ubuntu 24.04.5 LTS

Downloads

Two Linux x86-64 CUDA flavours. Pick by GPU generation and installed driver; both carry the same code.

asset CUDA GPUs with native code (SASS) PTX for other GPUs (JIT by the driver) minimum driver validation
llama.cpp-multigpu-20261006.1-bin-ubuntu-cuda-12.9-x64.tar.gz 12.9.1 sm_70, sm_86, sm_89, sm_120a, sm_121a sm_50, sm_61, sm_70, sm_75, sm_80, sm_90 R525+ (native targets); R570+ for Blackwell; R575+ for the PTX-only GPUs (CUDA 12.9 PTX JIT) Tier A passed; Tier B (sm_86, 6 x RTX 3090) by the maintainer
llama.cpp-multigpu-20261006.1-cudart-ubuntu-cuda-12.9-x64.tar.gz 12.9.1 runtime libraries only - R525+ (native targets); R570+ for Blackwell; R575+ for the PTX-only GPUs (CUDA 12.9 PTX JIT) -
llama.cpp-multigpu-20261006.1-bin-ubuntu-cuda-13.4-x64.tar.gz 13.4.1 sm_86, sm_89, sm_120a, sm_121a sm_80, sm_90 R580+ (CUDA 13); a driver as new as CUDA 13.4 for the PTX-only GPUs Tier A passed; not runtime-validated by us
llama.cpp-multigpu-20261006.1-cudart-ubuntu-cuda-13.4-x64.tar.gz 13.4.1 runtime libraries only - R580+ (CUDA 13); a driver as new as CUDA 13.4 for the PTX-only GPUs -
llama.cpp-multigpu-20261006.1-patches.tar.gz - git format-patch series - - -
  • cuda-12.9: the wide build. Native code for Volta (V100, sm_70), Ampere (sm_86), Ada (sm_89) and Blackwell
    (sm_120a/sm_121a); PTX for Maxwell, Pascal, Turing, A100 (sm_80) and Hopper (sm_90), which the driver compiles
    at first load. Needs an R525+ driver for the native targets, R570+ for Blackwell cards, and a driver that understands
    CUDA 12.9 PTX (R575+) for the PTX-only GPUs.

  • cuda-13.4: Ampere and newer only (CUDA 13 dropped Maxwell/Pascal/Volta compilation). Native sm_86/sm_89/
    sm_120a/sm_121a, PTX for sm_80/sm_90. Needs an R580+ driver.

  • cudart-*: the CUDA runtime + cuBLAS libraries matching each flavour, for hosts without a CUDA toolkit. Unpack next to
    the binaries (or set LD_LIBRARY_PATH). The NVIDIA driver (libcuda.so.1) always comes from the host.

  • patches.tar.gz: the patch series as git format-patch output, for building upstream + selected patches.

  • Container images (public, anonymous pull; built from these very archives; llama-server is the entrypoint, run
    with --gpus all) — package page: https://github.com/lukolszewski/llama.cpp-multigpu/pkgs/container/llama.cpp-multigpu

    image contents
    docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:server-cuda12.9-20261006.1 cuda-12.9 archive on nvidia/cuda:12.9.1-runtime-ubuntu24.04; V100 → RTX 5090; moving tags :server-cuda12.9, :latest
    docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:server-cuda13.4-20261006.1 cuda-13.4 archive on nvidia/cuda:13.4.1-runtime-ubuntu24.04; Ampere and newer; moving tag :server-cuda13.4; the base image requires a driver reporting CUDA ≥ 13.4 (-e NVIDIA_DISABLE_REQUIRE=1 on 13.x drivers)
    docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:bench-cuda12.9-20261006.1 the cuda-12.9 server image plus the README benchmark grid runner for rented machines (scripts/multigpu/bench/vast/README.md); moving tag :bench-cuda12.9

    Upstream baseline used for the comparisons: https://github.com/lukolszewski/llama.cpp-multigpu/releases/tag/baseline-df03399
    (unmodified upstream df03399b8, same cuda-12.9 recipe, tarball only).

Validation: Tier A = the offline packaging gate passed in CI (archive unpacks, binaries run, the declared SASS/PTX
targets are really embedded, the downstream flags are wired, metadata matches the commit). Tier B ...

Read more

Upstream baseline df03399b8 (b10902), cuda-12.9 recipe

Choose a tag to compare

@github-actions github-actions released this 06 Oct 17:58

Unmodified upstream llama.cpp commit df03399b8 (build b10902), compiled with the cuda-12.9 recipe of this repository's release workflow (CUDA 12.9.1, SASS sm_70/86/89/120a/121a, PTX sm_50/61/70/75/80/90), no downstream patches.

It is the upstream side of the benchmark grids in benches/multi-gpu/ and is fetched by the bench image (scripts/multigpu/bench/vast/). Not a release of llama.cpp-multigpu; for those see the multigpu-* releases. Needs the CUDA 12.9 runtime libraries (cudart-… archive of any multigpu-* release, or NVIDIA's nvidia/cuda:12.9.1-runtime image) and an R525+ driver (R575+ for PTX-only GPUs).