Skip to content

llama.cpp-multigpu 20261006.1 — Qwen3.8-Flash-Next multi-GPU performance build

Latest

Choose a tag to compare

@github-actions github-actions released this 06 Oct 19:23
· 10 commits to multigpu since this release
79a12ee

llama.cpp-multigpu

Upstream llama.cpp plus 47 performance patches for Qwen3.8-Flash-Next (qwen4exp) on consumer multi-GPU systems with
constrained PCIe (measured on 6 × RTX 3090). Same engine, same GGUF models, same llama-server API; the whole patchset is
selected at runtime by environment variables and a few server flags, so one build carries all of it. This is the
configuration running in production on the benchmark machine since 2026-10-05.

Measured against upstream (same machine, same command line; upstream = the commit this patchset branches from)

workload 5k 50k 150k 200k 250k tokens
1 slot prefill, upstream → multigpu (t/s) 715 → 1249 568 → 2127 368 → 2152 315 → 2122 263 → 2111
1 slot decode (t/s) 38.5 → 45.9 25.7 → 41.9 14.1 → 37.8 11.9 → 36.6 10.2 → 33.7
5 slots prefill, aggregate (t/s) 630 → 2120 574 → 2301 363 → 2251 307 → 2180 265 → 2060
5 slots decode, per slot (t/s) 18.2 → 42.3 8.1 → 38.3 3.6 → 32.4 2.8 → 29.0 2.3 → 27.3

Protocol, configuration and raw data: docs/multigpu/benchmarks.md, benches/multi-gpu/machine-01-7950x-6x3090/.

What the patches do (every switch and commit: docs/multigpu/patches.md)

  1. Long-context decode and quantized KV (patches 1–2). Fixes the CUDA decode slowdown of qwen4exp at depth and adds
    a compact-gather attention path so a q8_0 KV cache keeps the sparse-attention speedup (upstream falls back to a dense
    path): 225k-context decode 22 → 44 t/s.
  2. Server scheduling (3–4). Busy slots stay on contiguous sequence ids so several sessions decode in one ubatch; a
    prefill admission policy bounds how many prompts prefill at once (--prefill-max-partial, --prefill-max-long).
  3. Pipeline parallelism and device-built inputs (5–8). Ubatches overlap across GPUs for this hybrid model
    (LLAMA_PIPELINE_PARALLEL=1, GGML_CUDA_GRAPHS_FORCE=1); the causal KQ mask and the QSA block bias are built on the
    device instead of being copied over PCIe every ubatch: prefill at 82k 1260 → 2360 t/s, flat to 250k.
  4. Per-sequence decode pipeline and decode groups (9–27). LLAMA_DECODE_PIPELINE=1 splits multi-user decode into
    per-sequence ubatches on a stream-agnostic graph; LLAMA_SERVER_GROUPS=5 keeps several groups' batches in flight so
    users' tokens overlap across the GPUs instead of waiting for each other. 5 users at 58k: 17 → 34 t/s each.
  5. Prefill while others decode (28–36). Prompts are prefilled in 256-token chunks beside the decoding slots without
    draining the pipeline: two users decoding during a 200k prefill see ~280 ms between tokens instead of 1.3 s.
  6. Asynchronous prompt cache (37–43). Slot save/restore of multi-GiB KV states runs on worker threads with pinned
    staging and fences; context checkpoints no longer drain the pipeline; a 2.4 GiB restore takes ~1 s while the other
    users keep decoding (worst hitch 0.4 s).
  7. n-gram lookup speculation with the groups (44–47). Prompt-lookup decoding (no draft model) made to work with the
    decode groups through recurrent-state snapshots in the memory, with a cap on drafting users and an acceptance gate.
    Off by default; --spec-type ngram-map-k4v + LLAMA_SERVER_SPEC_MAX_USERS=2: solo code rewrite 44 → 111 t/s.

Also fixed on the way: an output-slot leak that crashed the server after client disconnects under load, an output-buffer
reallocation that corrupted parked results, and the upstream n-gram map keeping stale state across prompts.

Reference configuration (the numbers above were taken with it)

LLAMA_DECODE_PIPELINE=1 LLAMA_SERVER_GROUPS=5 LLAMA_PIPELINE_PARALLEL=1 GGML_CUDA_GRAPHS_FORCE=1 LLAMA_ATTN_ROT_DISABLE=1 \
llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -ngl 99 -fa on -np 5 -c 1310720 \
  --cache-type-k q8_0 --cache-type-v q8_0 -b 2048 -ub 512 -ot 'per_layer_token_embd\.weight=CPU' \
  --tensor-split 0.85,1,1,1,1,1 --prefill-max-partial 2

Scope and caveats

Targeted at one model family on one class of hardware; other models and NVLink or single-GPU machines may not benefit and
can regress. Only sm_86 is runtime-validated by the maintainer; the other architectures in the archives are compiled, not
run. Mixed prefill + decode is better than upstream but still a documented limitation. Upstream llama.cpp remains the
general-purpose project; this fork exists to be archived once upstream catches up.

Container images: forward-compatibility library removed (from multigpu-20261006.1)

NVIDIA's cuda:*-runtime base images ship cuda-compat-<ver>, a newer libcuda for datacenter GPUs on old
drivers. On a host whose driver is older than that library the container toolkit uses it, GeForce cards reject it
(ggml_cuda_init: failed to initialize CUDA: forward compatibility was attempted on non supported HW) and
llama.cpp silently runs on the CPU. Found on an 8 × RTX 4090 host with driver 570 using the multigpu-20261006
images; the package is purged from every image from this release on, so the host driver's libcuda is used
(any R525+ driver for the native targets). If you must run the multigpu-20261006 images on such a host:
-v /dev/null:/usr/local/cuda/compat/libcuda.so.1. The tarballs were never affected. Details:
docs/multigpu/builds.md.

This build

release tag multigpu-20261006.1
multigpu commit 79a12ee4767d (branch multigpu)
upstream llama.cpp base df03399b8858 (2026-09-10)
commits ahead of upstream 94 (47 code patches indexed in docs/multigpu/patches.md, the rest docs/CI; full list below, also attached as -patches.tar.gz)
toolchain cuda-12.9 CUDA 12.9.1 (release 12.9), cc (Ubuntu 13.3.0-6ubuntu2~24.04) 13.3.0, cmake version 3.28.3, Ubuntu 24.04.2 LTS
toolchain cuda-13.4 CUDA 13.4.1 (release 13.4), cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0, cmake version 3.28.3, Ubuntu 24.04.5 LTS

Downloads

Two Linux x86-64 CUDA flavours. Pick by GPU generation and installed driver; both carry the same code.

asset CUDA GPUs with native code (SASS) PTX for other GPUs (JIT by the driver) minimum driver validation
llama.cpp-multigpu-20261006.1-bin-ubuntu-cuda-12.9-x64.tar.gz 12.9.1 sm_70, sm_86, sm_89, sm_120a, sm_121a sm_50, sm_61, sm_70, sm_75, sm_80, sm_90 R525+ (native targets); R570+ for Blackwell; R575+ for the PTX-only GPUs (CUDA 12.9 PTX JIT) Tier A passed; Tier B (sm_86, 6 x RTX 3090) by the maintainer
llama.cpp-multigpu-20261006.1-cudart-ubuntu-cuda-12.9-x64.tar.gz 12.9.1 runtime libraries only - R525+ (native targets); R570+ for Blackwell; R575+ for the PTX-only GPUs (CUDA 12.9 PTX JIT) -
llama.cpp-multigpu-20261006.1-bin-ubuntu-cuda-13.4-x64.tar.gz 13.4.1 sm_86, sm_89, sm_120a, sm_121a sm_80, sm_90 R580+ (CUDA 13); a driver as new as CUDA 13.4 for the PTX-only GPUs Tier A passed; not runtime-validated by us
llama.cpp-multigpu-20261006.1-cudart-ubuntu-cuda-13.4-x64.tar.gz 13.4.1 runtime libraries only - R580+ (CUDA 13); a driver as new as CUDA 13.4 for the PTX-only GPUs -
llama.cpp-multigpu-20261006.1-patches.tar.gz - git format-patch series - - -
  • cuda-12.9: the wide build. Native code for Volta (V100, sm_70), Ampere (sm_86), Ada (sm_89) and Blackwell
    (sm_120a/sm_121a); PTX for Maxwell, Pascal, Turing, A100 (sm_80) and Hopper (sm_90), which the driver compiles
    at first load. Needs an R525+ driver for the native targets, R570+ for Blackwell cards, and a driver that understands
    CUDA 12.9 PTX (R575+) for the PTX-only GPUs.

  • cuda-13.4: Ampere and newer only (CUDA 13 dropped Maxwell/Pascal/Volta compilation). Native sm_86/sm_89/
    sm_120a/sm_121a, PTX for sm_80/sm_90. Needs an R580+ driver.

  • cudart-*: the CUDA runtime + cuBLAS libraries matching each flavour, for hosts without a CUDA toolkit. Unpack next to
    the binaries (or set LD_LIBRARY_PATH). The NVIDIA driver (libcuda.so.1) always comes from the host.

  • patches.tar.gz: the patch series as git format-patch output, for building upstream + selected patches.

  • Container images (public, anonymous pull; built from these very archives; llama-server is the entrypoint, run
    with --gpus all) — package page: https://github.com/lukolszewski/llama.cpp-multigpu/pkgs/container/llama.cpp-multigpu

    image contents
    docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:server-cuda12.9-20261006.1 cuda-12.9 archive on nvidia/cuda:12.9.1-runtime-ubuntu24.04; V100 → RTX 5090; moving tags :server-cuda12.9, :latest
    docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:server-cuda13.4-20261006.1 cuda-13.4 archive on nvidia/cuda:13.4.1-runtime-ubuntu24.04; Ampere and newer; moving tag :server-cuda13.4; the base image requires a driver reporting CUDA ≥ 13.4 (-e NVIDIA_DISABLE_REQUIRE=1 on 13.x drivers)
    docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:bench-cuda12.9-20261006.1 the cuda-12.9 server image plus the README benchmark grid runner for rented machines (scripts/multigpu/bench/vast/README.md); moving tag :bench-cuda12.9

    Upstream baseline used for the comparisons: https://github.com/lukolszewski/llama.cpp-multigpu/releases/tag/baseline-df03399
    (unmodified upstream df03399b8, same cuda-12.9 recipe, tarball only).

Validation: Tier A = the offline packaging gate passed in CI (archive unpacks, binaries run, the declared SASS/PTX
targets are really embedded, the downstream flags are wired, metadata matches the commit). Tier B = served a model on
real hardware; only Linux / sm_86 (6 x RTX 3090) is run by the maintainer. Every other GPU architecture in these archives
is compiled, not run - reports from such hardware are welcome as issues (archive name, GPU model/count, driver).
SHA256SUMS.txt covers every asset; each archive embeds BUILD_INFO.json / BUILD_INFO.txt.

All 94 commits ahead of upstream in this build (code patches and docs/CI) - per-patch description and switches: docs/multigpu/patches.md
commit subject
b22bc5b68 qwen4exp: fix CUDA decode depth-decay at long context (re: ggml-org#28734)
4d6ef5ad4 qwen4exp: compact-gather attention for decode — quantized KV keeps the sparse speedup (re: ggml-org#28734)
b75e380e0 server: keep busy slots on contiguous sequence ids (seq compaction)
02e136bef server: prefill admission policy (max-partial / long-threshold / max-long) + waiting metrics
c8832a5f0 pipeline parallelism for the qwen4exp hybrid: keep ubatches overlapping across GPUs
0f5aaf45d ggml-backend: GGML_SCHED_N_COPIES env to limit pipeline input copies (memory vs depth)
a4e856f1a graph: build the causal KQ mask on the device from cell positions (LLAMA_KQ_MASK_DEVICE)
9f5a3202e qwen4exp: build the QSA block bias on the device (LLAMA_QSA_BIAS_DEVICE)
4660a295e memory-hybrid-idx: LLAMA_DECODE_PIPELINE=1 splits a pure multi-sequence decode batch into per-sequence ubatches
fb8897947 stream-agnostic decode graph for per-sequence decode ubatches (LLAMA_DECODE_PIPELINE=1)
1567041ef LLAMA_UBATCH_TRACE=1: per-ubatch reuse/reserve/timing and per-split host-wait trace (debug aid)
1ca80a510 ggml-backend: stage scheduler inputs through rotating pinned buffers and copy them asynchronously on the compute stream
5110faa51 ggml-backend: one staging offset per backend per compute, not per split
6b7bb59fe ggml-backend: grow the staging slots geometrically, and a GGML_SCHED_STAGE_WAIT bisection knob
850db8aec ggml-backend: the source stream of a cross-backend hop copy waits for the destination's previous split (reused-graph pipelining fix; GGML_SCHED_NO_HOP_WAIT=1 disables)
21910bb3b ggml-backend: stage scheduler inputs only for reused graphs (a freshly allocated graph has a fresh copy slot); GGML_SCHED_STAGE_INPUTS=2 stages always
66a32a7f8 scheduler input staging only for pipelined decode ubatches (ggml_backend_sched_set_stage_inputs, set per ubatch by llama_context); solo decode and prefill use the plain synchronous copies
fa532f399 server: LLAMA_SERVER_TIMINGS=1 enables the per-phase update_slots timing report at runtime (5 s windows, reset after each report)
53fcc9f87 phase 2: multiple in-flight batches (llama_output_slots_set/llama_output_slot/llama_output_select: per-batch output regions and completion events) and server pipelined decode groups (LLAMA_SERVER_GROUPS=n): one group sampled and re-issued per update_slots while the others compute
d59899454 server: member-wise server_batch::swap for parking pipelined-group batches (std::swap double-freed the llama_batch arrays)
9109ac64c LLAMA_DECODE_PIPELINE=2: single-sequence decode batches also take the stream-agnostic class with a cache-wide n_kv floor (one graph for all pipelined-group batches); the server sets it when groups are on
c89c7fc89 trace: mem_hybrid reuse reject detail
ec1b7bf52 decode pipeline level read live from the environment (the static was captured by the warmup decode before the server set level 2)
2c9807e15 trace full synchronize while output slots are in use; timing report must not synchronize in group mode
00b3703f0 backend-sampling getters wait only for the selected output batch (the common sampler probes them on every token; the full synchronize serialized pipelined groups)
f8c8a9ce0 common sampler waits only for the selected output batch (llama_output_synchronize) instead of a full synchronize per token
ef7ffc7e2 server groups: skip groups without work, per-stream graph class while a single slot is active
07728cbfa server groups: chunk prefill to one ubatch and one slot while other slots generate
da3fbb65d context: reserve compute buffers once per ubatch shape class
6ccace4e8 sched: do not drain the pipeline when the graph layout changes
7419cd87c server groups: eager prefill chunks; sched: pre-allocated 12-slot staging ring
442b4e349 sched: synchronize before an unstaged host copy after a re-plan; server: eager prefill off by default
270b27d53 safe defaults: scheduler drain on re-plan, prefill chunking off
c73c59364 sched: cover the user's asynchronous output read-outs with the completion events; non-draining re-plan and 256-token prefill chunks by default
ca0faad46 sched: keep the re-plan safety window open for several computes
93764d163 sched: order re-planned computes after other-plan computes only; server: chunk only while decoders outnumber prefills
a1405f2f7 WIP (not compiled): asynchronous prompt-cache save/restore
fcb21fe5f server: drain-free context checkpoints, pinned staging for state transfers, idle saves spread out
1b85f6429 server: finish asynchronous state transfers even when no slot is processing
2866c72a2 server: a task whose best-matching slot is being saved waits for it
c8bef79c2 server: cheap prompt-cache save start (shared checkpoint bytes, off-thread frees)
4fb700cf5 llama: staged state copies on several threads (LLAMA_STATE_XFER_THREADS)
6a33bc322 server: release a parked batch's output slot even when its logits were never read
efef36982 server: n-gram speculation with decode groups, recurrent-state snapshots instead of host checkpoints
83f09b865 speculative: fresh n-gram map per new prompt; server: cap on the number of drafting users
7c29e21b8 server: acceptance gate for speculative drafts
e4054f726 llama: portable std::max in output_reserve (gcc 13 rejects the explicit-template brace-list form)
c37a0749e docs(multigpu): fork README with temporary-downstream framing
b35fa2285 docs(multigpu): patch index, benchmark protocol, build policy, upstream sync
d8af4f398 tools(multigpu): provenance metadata, patch export, Tier A gate, release notes
c876c5ac9 tools(multigpu): benchmark drivers, Tier B smoke and hardware snapshot
928bee326 tools(multigpu): upstream sync helper and fork workflow management
1cb99a9a1 ci(multigpu): build the patched branch, gate artifacts, export patches
3e70caf2d ci(multigpu): nightly rolling release plus date-tagged builds
bb1111b55 ci(multigpu): opt-in self-hosted Tier B validation and benchmark jobs
db7ab4c32 gh: templates for multi-GPU performance reports and benchmark submissions
3b7747660 ci: stop the upstream release pipeline from firing on this fork
134489582 docs(multigpu): patch index, usage and status for the 47-commit production patchset
b724ead7f benchmarks(multigpu): machine-01 grid, upstream df03399 vs multigpu (2026-10-05)
dc7f1eabf multigpu: repository renamed to lukolszewski/llama.cpp-multigpu
f7ec0f334 Update README.md
5a15c64f0 README: the summary row and limitations now reflect the measured benchmarks
81f62c01b docs(multigpu): release description intro (patch overview, measured numbers, reference configuration)
ea1878381 ci(multigpu): packaging-only runtime Dockerfile for release and CI archives
ade11b53e ci(multigpu): gate verifies PTX targets too; build metadata records SASS/PTX/driver floor
27af49fcc ci(multigpu): push builds one sm_86 image; tags build two CUDA flavours and the release
258b37ea4 docs(multigpu): builds, downloads and patch index for the two-flavour release scheme
a1f5ddf29 tests: load the ggml backends in test-llama-archs so test-generate-models works with GGML_BACKEND_DL
adf4bd7a0 ci(multigpu): install git-lfs in the push build for test-tokenizers-ggml-vocabs
6a8a5995f ci(multigpu): install git before checkout in the container jobs; ccache dir via GITHUB_WORKSPACE
8a772e3a6 docs(multigpu): the ghcr package is public already; how to check it
ab9e7ecbb README: lead with the measured speedup (TL;DR table + chart), performance first, badges
3c9be8cd9 ci(multigpu): benchmark tooling changes do not trigger the push build [skip ci]
022c3361f docs(multigpu): Tier B record for multigpu-20261006; the cuda-13.4 image needs a CUDA>=13.4 driver (or NVIDIA_DISABLE_REQUIRE=1)
c316e9ea0 bench: image, entrypoint and CI for benchmarking on rented multi-GPU machines
b8c7da00e bench: vast-bench.sh orchestrator (search/run/status/fetch/land/destroy); baseline workflow uses bash
7aee6b309 bench: vast docs, test-local.sh, README/benchmarks/builds sections; orchestrator fixes from the rehearsal
3be5ee84e bench: vast-bench.sh --ingress-gb for the cost estimate
defcf4ba0 bench: vast-bench.sh state file: jq 1.6 rejects $label (keyword)
5b66a996c validate-artifact: parse the commit from upstream's '(build N, commit SHA)' --version format
c916edd21 bench: vast-bench.sh: --cancel-unavail, detect stopped instances, ssh/log diagnostics on failure
345900d92 bench image: sshd StrictModes no (Vast's authorized_keys modes), key-only root login; orchestrator ssh IdentitiesOnly=yes
ed042690e bench: vast README: findings from the rehearsals (destroy -y, stopped instances, sshd StrictModes, execute limits)
06929d2c3 bench: config.json records every knob (export before the dump); land drops the marker dotfiles
386388e5e bench: gh pr create names the repo (gh targets the fork parent otherwise); image workflow only rebuilds on files that are in the image
a0deb79e6 bench: entrypoint logs the download directly (the tail -n 0 pipeline SIGPIPEd aria2: exit 141 on machine-02); aria2 readout off in logs
dc5c7a8d5 bench: download progress counts allocated bytes (sparse aria2 files made the stat-based monitor report 126 %)
a5178e3e9 baseline: gate checks only fork-only flags (--cache-idle-slots exists upstream too)
0c0d41ca2 bench: purge cuda-compat from both images (GeForce + older driver -> CPU fallback); pre-flight asks llama-server --list-devices and checks VRAM after load; orchestrator aborts on CPU-speed rows and keeps the box for in-place repair
076c78865 bench: wait up to UPSTREAM_WAIT (1 h) for the baseline tarball instead of failing
a013d01eb bench: vast-bench.sh --machine-id (A/B on the same physical box)
c952cfdd3 bench: land picks a free branch name when a run of the same machine landed the same day
ea72de808 docs: release intro notes the removed forward-compatibility library
79a12ee47 bench: benchmark grid on rented multi-GPU machines (Vast.ai); images without cuda-compat (#1)

License: MIT (upstream llama.cpp license preserved). Upstream llama.cpp is where this work belongs; this fork exists to be archived once upstream performs comparably.