Repository navigation
llama.cpp-multigpu
Upstream llama.cpp plus 47 performance patches for Qwen3.8-Flash-Next (qwen4exp) on consumer multi-GPU systems with
constrained PCIe (measured on 6 × RTX 3090). Same engine, same GGUF models, same llama-server API; the whole patchset is
selected at runtime by environment variables and a few server flags, so one build carries all of it. This is the
configuration running in production on the benchmark machine since 2026-10-05.
Measured against upstream (same machine, same command line; upstream = the commit this patchset branches from)
| workload | 5k | 50k | 150k | 200k | 250k tokens |
|---|---|---|---|---|---|
| 1 slot prefill, upstream → multigpu (t/s) | 715 → 1249 | 568 → 2127 | 368 → 2152 | 315 → 2122 | 263 → 2111 |
| 1 slot decode (t/s) | 38.5 → 45.9 | 25.7 → 41.9 | 14.1 → 37.8 | 11.9 → 36.6 | 10.2 → 33.7 |
| 5 slots prefill, aggregate (t/s) | 630 → 2120 | 574 → 2301 | 363 → 2251 | 307 → 2180 | 265 → 2060 |
| 5 slots decode, per slot (t/s) | 18.2 → 42.3 | 8.1 → 38.3 | 3.6 → 32.4 | 2.8 → 29.0 | 2.3 → 27.3 |
Protocol, configuration and raw data: docs/multigpu/benchmarks.md, benches/multi-gpu/machine-01-7950x-6x3090/.
What the patches do (every switch and commit: docs/multigpu/patches.md)
- Long-context decode and quantized KV (patches 1–2). Fixes the CUDA decode slowdown of
qwen4expat depth and adds
a compact-gather attention path so aq8_0KV cache keeps the sparse-attention speedup (upstream falls back to a dense
path): 225k-context decode 22 → 44 t/s. - Server scheduling (3–4). Busy slots stay on contiguous sequence ids so several sessions decode in one ubatch; a
prefill admission policy bounds how many prompts prefill at once (--prefill-max-partial,--prefill-max-long). - Pipeline parallelism and device-built inputs (5–8). Ubatches overlap across GPUs for this hybrid model
(LLAMA_PIPELINE_PARALLEL=1,GGML_CUDA_GRAPHS_FORCE=1); the causal KQ mask and the QSA block bias are built on the
device instead of being copied over PCIe every ubatch: prefill at 82k 1260 → 2360 t/s, flat to 250k. - Per-sequence decode pipeline and decode groups (9–27).
LLAMA_DECODE_PIPELINE=1splits multi-user decode into
per-sequence ubatches on a stream-agnostic graph;LLAMA_SERVER_GROUPS=5keeps several groups' batches in flight so
users' tokens overlap across the GPUs instead of waiting for each other. 5 users at 58k: 17 → 34 t/s each. - Prefill while others decode (28–36). Prompts are prefilled in 256-token chunks beside the decoding slots without
draining the pipeline: two users decoding during a 200k prefill see ~280 ms between tokens instead of 1.3 s. - Asynchronous prompt cache (37–43). Slot save/restore of multi-GiB KV states runs on worker threads with pinned
staging and fences; context checkpoints no longer drain the pipeline; a 2.4 GiB restore takes ~1 s while the other
users keep decoding (worst hitch 0.4 s). - n-gram lookup speculation with the groups (44–47). Prompt-lookup decoding (no draft model) made to work with the
decode groups through recurrent-state snapshots in the memory, with a cap on drafting users and an acceptance gate.
Off by default;--spec-type ngram-map-k4v+LLAMA_SERVER_SPEC_MAX_USERS=2: solo code rewrite 44 → 111 t/s.
Also fixed on the way: an output-slot leak that crashed the server after client disconnects under load, an output-buffer
reallocation that corrupted parked results, and the upstream n-gram map keeping stale state across prompts.
Reference configuration (the numbers above were taken with it)
LLAMA_DECODE_PIPELINE=1 LLAMA_SERVER_GROUPS=5 LLAMA_PIPELINE_PARALLEL=1 GGML_CUDA_GRAPHS_FORCE=1 LLAMA_ATTN_ROT_DISABLE=1 \
llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -ngl 99 -fa on -np 5 -c 1310720 \
--cache-type-k q8_0 --cache-type-v q8_0 -b 2048 -ub 512 -ot 'per_layer_token_embd\.weight=CPU' \
--tensor-split 0.85,1,1,1,1,1 --prefill-max-partial 2
Scope and caveats
Targeted at one model family on one class of hardware; other models and NVLink or single-GPU machines may not benefit and
can regress. Only sm_86 is runtime-validated by the maintainer; the other architectures in the archives are compiled, not
run. Mixed prefill + decode is better than upstream but still a documented limitation. Upstream llama.cpp remains the
general-purpose project; this fork exists to be archived once upstream catches up.
Container images: forward-compatibility library removed (from multigpu-20261006.1)
NVIDIA's cuda:*-runtime base images ship cuda-compat-<ver>, a newer libcuda for datacenter GPUs on old
drivers. On a host whose driver is older than that library the container toolkit uses it, GeForce cards reject it
(ggml_cuda_init: failed to initialize CUDA: forward compatibility was attempted on non supported HW) and
llama.cpp silently runs on the CPU. Found on an 8 × RTX 4090 host with driver 570 using the multigpu-20261006
images; the package is purged from every image from this release on, so the host driver's libcuda is used
(any R525+ driver for the native targets). If you must run the multigpu-20261006 images on such a host:
-v /dev/null:/usr/local/cuda/compat/libcuda.so.1. The tarballs were never affected. Details:
docs/multigpu/builds.md.
This build
| release tag | multigpu-20261006.1 |
| multigpu commit | 79a12ee4767d (branch multigpu) |
| upstream llama.cpp base | df03399b8858 (2026-09-10) |
| commits ahead of upstream | 94 (47 code patches indexed in docs/multigpu/patches.md, the rest docs/CI; full list below, also attached as -patches.tar.gz) |
toolchain cuda-12.9 |
CUDA 12.9.1 (release 12.9), cc (Ubuntu 13.3.0-6ubuntu2~24.04) 13.3.0, cmake version 3.28.3, Ubuntu 24.04.2 LTS |
toolchain cuda-13.4 |
CUDA 13.4.1 (release 13.4), cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0, cmake version 3.28.3, Ubuntu 24.04.5 LTS |
Downloads
Two Linux x86-64 CUDA flavours. Pick by GPU generation and installed driver; both carry the same code.
| asset | CUDA | GPUs with native code (SASS) | PTX for other GPUs (JIT by the driver) | minimum driver | validation |
|---|---|---|---|---|---|
llama.cpp-multigpu-20261006.1-bin-ubuntu-cuda-12.9-x64.tar.gz |
12.9.1 | sm_70, sm_86, sm_89, sm_120a, sm_121a | sm_50, sm_61, sm_70, sm_75, sm_80, sm_90 | R525+ (native targets); R570+ for Blackwell; R575+ for the PTX-only GPUs (CUDA 12.9 PTX JIT) | Tier A passed; Tier B (sm_86, 6 x RTX 3090) by the maintainer |
llama.cpp-multigpu-20261006.1-cudart-ubuntu-cuda-12.9-x64.tar.gz |
12.9.1 | runtime libraries only | - | R525+ (native targets); R570+ for Blackwell; R575+ for the PTX-only GPUs (CUDA 12.9 PTX JIT) | - |
llama.cpp-multigpu-20261006.1-bin-ubuntu-cuda-13.4-x64.tar.gz |
13.4.1 | sm_86, sm_89, sm_120a, sm_121a | sm_80, sm_90 | R580+ (CUDA 13); a driver as new as CUDA 13.4 for the PTX-only GPUs | Tier A passed; not runtime-validated by us |
llama.cpp-multigpu-20261006.1-cudart-ubuntu-cuda-13.4-x64.tar.gz |
13.4.1 | runtime libraries only | - | R580+ (CUDA 13); a driver as new as CUDA 13.4 for the PTX-only GPUs | - |
llama.cpp-multigpu-20261006.1-patches.tar.gz |
- | git format-patch series | - | - | - |
-
cuda-12.9: the wide build. Native code for Volta (V100,sm_70), Ampere (sm_86), Ada (sm_89) and Blackwell
(sm_120a/sm_121a); PTX for Maxwell, Pascal, Turing, A100 (sm_80) and Hopper (sm_90), which the driver compiles
at first load. Needs an R525+ driver for the native targets, R570+ for Blackwell cards, and a driver that understands
CUDA 12.9 PTX (R575+) for the PTX-only GPUs. -
cuda-13.4: Ampere and newer only (CUDA 13 dropped Maxwell/Pascal/Volta compilation). Nativesm_86/sm_89/
sm_120a/sm_121a, PTX forsm_80/sm_90. Needs an R580+ driver. -
cudart-*: the CUDA runtime + cuBLAS libraries matching each flavour, for hosts without a CUDA toolkit. Unpack next to
the binaries (or setLD_LIBRARY_PATH). The NVIDIA driver (libcuda.so.1) always comes from the host. -
patches.tar.gz: the patch series asgit format-patchoutput, for building upstream + selected patches. -
Container images (public, anonymous pull; built from these very archives;
llama-serveris the entrypoint, run
with--gpus all) — package page: https://github.com/lukolszewski/llama.cpp-multigpu/pkgs/container/llama.cpp-multigpuimage contents docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:server-cuda12.9-20261006.1cuda-12.9 archive on nvidia/cuda:12.9.1-runtime-ubuntu24.04; V100 → RTX 5090; moving tags:server-cuda12.9,:latestdocker pull ghcr.io/lukolszewski/llama.cpp-multigpu:server-cuda13.4-20261006.1cuda-13.4 archive on nvidia/cuda:13.4.1-runtime-ubuntu24.04; Ampere and newer; moving tag:server-cuda13.4; the base image requires a driver reporting CUDA ≥ 13.4 (-e NVIDIA_DISABLE_REQUIRE=1on 13.x drivers)docker pull ghcr.io/lukolszewski/llama.cpp-multigpu:bench-cuda12.9-20261006.1the cuda-12.9 server image plus the README benchmark grid runner for rented machines ( scripts/multigpu/bench/vast/README.md); moving tag:bench-cuda12.9Upstream baseline used for the comparisons: https://github.com/lukolszewski/llama.cpp-multigpu/releases/tag/baseline-df03399
(unmodified upstreamdf03399b8, same cuda-12.9 recipe, tarball only).
Validation: Tier A = the offline packaging gate passed in CI (archive unpacks, binaries run, the declared SASS/PTX
targets are really embedded, the downstream flags are wired, metadata matches the commit). Tier B = served a model on
real hardware; only Linux / sm_86 (6 x RTX 3090) is run by the maintainer. Every other GPU architecture in these archives
is compiled, not run - reports from such hardware are welcome as issues (archive name, GPU model/count, driver).
SHA256SUMS.txt covers every asset; each archive embeds BUILD_INFO.json / BUILD_INFO.txt.
All 94 commits ahead of upstream in this build (code patches and docs/CI) - per-patch description and switches: docs/multigpu/patches.md
| commit | subject |
|---|---|
b22bc5b68 |
qwen4exp: fix CUDA decode depth-decay at long context (re: ggml-org#28734) |
4d6ef5ad4 |
qwen4exp: compact-gather attention for decode — quantized KV keeps the sparse speedup (re: ggml-org#28734) |
b75e380e0 |
server: keep busy slots on contiguous sequence ids (seq compaction) |
02e136bef |
server: prefill admission policy (max-partial / long-threshold / max-long) + waiting metrics |
c8832a5f0 |
pipeline parallelism for the qwen4exp hybrid: keep ubatches overlapping across GPUs |
0f5aaf45d |
ggml-backend: GGML_SCHED_N_COPIES env to limit pipeline input copies (memory vs depth) |
a4e856f1a |
graph: build the causal KQ mask on the device from cell positions (LLAMA_KQ_MASK_DEVICE) |
9f5a3202e |
qwen4exp: build the QSA block bias on the device (LLAMA_QSA_BIAS_DEVICE) |
4660a295e |
memory-hybrid-idx: LLAMA_DECODE_PIPELINE=1 splits a pure multi-sequence decode batch into per-sequence ubatches |
fb8897947 |
stream-agnostic decode graph for per-sequence decode ubatches (LLAMA_DECODE_PIPELINE=1) |
1567041ef |
LLAMA_UBATCH_TRACE=1: per-ubatch reuse/reserve/timing and per-split host-wait trace (debug aid) |
1ca80a510 |
ggml-backend: stage scheduler inputs through rotating pinned buffers and copy them asynchronously on the compute stream |
5110faa51 |
ggml-backend: one staging offset per backend per compute, not per split |
6b7bb59fe |
ggml-backend: grow the staging slots geometrically, and a GGML_SCHED_STAGE_WAIT bisection knob |
850db8aec |
ggml-backend: the source stream of a cross-backend hop copy waits for the destination's previous split (reused-graph pipelining fix; GGML_SCHED_NO_HOP_WAIT=1 disables) |
21910bb3b |
ggml-backend: stage scheduler inputs only for reused graphs (a freshly allocated graph has a fresh copy slot); GGML_SCHED_STAGE_INPUTS=2 stages always |
66a32a7f8 |
scheduler input staging only for pipelined decode ubatches (ggml_backend_sched_set_stage_inputs, set per ubatch by llama_context); solo decode and prefill use the plain synchronous copies |
fa532f399 |
server: LLAMA_SERVER_TIMINGS=1 enables the per-phase update_slots timing report at runtime (5 s windows, reset after each report) |
53fcc9f87 |
phase 2: multiple in-flight batches (llama_output_slots_set/llama_output_slot/llama_output_select: per-batch output regions and completion events) and server pipelined decode groups (LLAMA_SERVER_GROUPS=n): one group sampled and re-issued per update_slots while the others compute |
d59899454 |
server: member-wise server_batch::swap for parking pipelined-group batches (std::swap double-freed the llama_batch arrays) |
9109ac64c |
LLAMA_DECODE_PIPELINE=2: single-sequence decode batches also take the stream-agnostic class with a cache-wide n_kv floor (one graph for all pipelined-group batches); the server sets it when groups are on |
c89c7fc89 |
trace: mem_hybrid reuse reject detail |
ec1b7bf52 |
decode pipeline level read live from the environment (the static was captured by the warmup decode before the server set level 2) |
2c9807e15 |
trace full synchronize while output slots are in use; timing report must not synchronize in group mode |
00b3703f0 |
backend-sampling getters wait only for the selected output batch (the common sampler probes them on every token; the full synchronize serialized pipelined groups) |
f8c8a9ce0 |
common sampler waits only for the selected output batch (llama_output_synchronize) instead of a full synchronize per token |
ef7ffc7e2 |
server groups: skip groups without work, per-stream graph class while a single slot is active |
07728cbfa |
server groups: chunk prefill to one ubatch and one slot while other slots generate |
da3fbb65d |
context: reserve compute buffers once per ubatch shape class |
6ccace4e8 |
sched: do not drain the pipeline when the graph layout changes |
7419cd87c |
server groups: eager prefill chunks; sched: pre-allocated 12-slot staging ring |
442b4e349 |
sched: synchronize before an unstaged host copy after a re-plan; server: eager prefill off by default |
270b27d53 |
safe defaults: scheduler drain on re-plan, prefill chunking off |
c73c59364 |
sched: cover the user's asynchronous output read-outs with the completion events; non-draining re-plan and 256-token prefill chunks by default |
ca0faad46 |
sched: keep the re-plan safety window open for several computes |
93764d163 |
sched: order re-planned computes after other-plan computes only; server: chunk only while decoders outnumber prefills |
a1405f2f7 |
WIP (not compiled): asynchronous prompt-cache save/restore |
fcb21fe5f |
server: drain-free context checkpoints, pinned staging for state transfers, idle saves spread out |
1b85f6429 |
server: finish asynchronous state transfers even when no slot is processing |
2866c72a2 |
server: a task whose best-matching slot is being saved waits for it |
c8bef79c2 |
server: cheap prompt-cache save start (shared checkpoint bytes, off-thread frees) |
4fb700cf5 |
llama: staged state copies on several threads (LLAMA_STATE_XFER_THREADS) |
6a33bc322 |
server: release a parked batch's output slot even when its logits were never read |
efef36982 |
server: n-gram speculation with decode groups, recurrent-state snapshots instead of host checkpoints |
83f09b865 |
speculative: fresh n-gram map per new prompt; server: cap on the number of drafting users |
7c29e21b8 |
server: acceptance gate for speculative drafts |
e4054f726 |
llama: portable std::max in output_reserve (gcc 13 rejects the explicit-template brace-list form) |
c37a0749e |
docs(multigpu): fork README with temporary-downstream framing |
b35fa2285 |
docs(multigpu): patch index, benchmark protocol, build policy, upstream sync |
d8af4f398 |
tools(multigpu): provenance metadata, patch export, Tier A gate, release notes |
c876c5ac9 |
tools(multigpu): benchmark drivers, Tier B smoke and hardware snapshot |
928bee326 |
tools(multigpu): upstream sync helper and fork workflow management |
1cb99a9a1 |
ci(multigpu): build the patched branch, gate artifacts, export patches |
3e70caf2d |
ci(multigpu): nightly rolling release plus date-tagged builds |
bb1111b55 |
ci(multigpu): opt-in self-hosted Tier B validation and benchmark jobs |
db7ab4c32 |
gh: templates for multi-GPU performance reports and benchmark submissions |
3b7747660 |
ci: stop the upstream release pipeline from firing on this fork |
134489582 |
docs(multigpu): patch index, usage and status for the 47-commit production patchset |
b724ead7f |
benchmarks(multigpu): machine-01 grid, upstream df03399 vs multigpu (2026-10-05) |
dc7f1eabf |
multigpu: repository renamed to lukolszewski/llama.cpp-multigpu |
f7ec0f334 |
Update README.md |
5a15c64f0 |
README: the summary row and limitations now reflect the measured benchmarks |
81f62c01b |
docs(multigpu): release description intro (patch overview, measured numbers, reference configuration) |
ea1878381 |
ci(multigpu): packaging-only runtime Dockerfile for release and CI archives |
ade11b53e |
ci(multigpu): gate verifies PTX targets too; build metadata records SASS/PTX/driver floor |
27af49fcc |
ci(multigpu): push builds one sm_86 image; tags build two CUDA flavours and the release |
258b37ea4 |
docs(multigpu): builds, downloads and patch index for the two-flavour release scheme |
a1f5ddf29 |
tests: load the ggml backends in test-llama-archs so test-generate-models works with GGML_BACKEND_DL |
adf4bd7a0 |
ci(multigpu): install git-lfs in the push build for test-tokenizers-ggml-vocabs |
6a8a5995f |
ci(multigpu): install git before checkout in the container jobs; ccache dir via GITHUB_WORKSPACE |
8a772e3a6 |
docs(multigpu): the ghcr package is public already; how to check it |
ab9e7ecbb |
README: lead with the measured speedup (TL;DR table + chart), performance first, badges |
3c9be8cd9 |
ci(multigpu): benchmark tooling changes do not trigger the push build [skip ci] |
022c3361f |
docs(multigpu): Tier B record for multigpu-20261006; the cuda-13.4 image needs a CUDA>=13.4 driver (or NVIDIA_DISABLE_REQUIRE=1) |
c316e9ea0 |
bench: image, entrypoint and CI for benchmarking on rented multi-GPU machines |
b8c7da00e |
bench: vast-bench.sh orchestrator (search/run/status/fetch/land/destroy); baseline workflow uses bash |
7aee6b309 |
bench: vast docs, test-local.sh, README/benchmarks/builds sections; orchestrator fixes from the rehearsal |
3be5ee84e |
bench: vast-bench.sh --ingress-gb for the cost estimate |
defcf4ba0 |
bench: vast-bench.sh state file: jq 1.6 rejects $label (keyword) |
5b66a996c |
validate-artifact: parse the commit from upstream's '(build N, commit SHA)' --version format |
c916edd21 |
bench: vast-bench.sh: --cancel-unavail, detect stopped instances, ssh/log diagnostics on failure |
345900d92 |
bench image: sshd StrictModes no (Vast's authorized_keys modes), key-only root login; orchestrator ssh IdentitiesOnly=yes |
ed042690e |
bench: vast README: findings from the rehearsals (destroy -y, stopped instances, sshd StrictModes, execute limits) |
06929d2c3 |
bench: config.json records every knob (export before the dump); land drops the marker dotfiles |
386388e5e |
bench: gh pr create names the repo (gh targets the fork parent otherwise); image workflow only rebuilds on files that are in the image |
a0deb79e6 |
bench: entrypoint logs the download directly (the tail -n 0 pipeline SIGPIPEd aria2: exit 141 on machine-02); aria2 readout off in logs |
dc5c7a8d5 |
bench: download progress counts allocated bytes (sparse aria2 files made the stat-based monitor report 126 %) |
a5178e3e9 |
baseline: gate checks only fork-only flags (--cache-idle-slots exists upstream too) |
0c0d41ca2 |
bench: purge cuda-compat from both images (GeForce + older driver -> CPU fallback); pre-flight asks llama-server --list-devices and checks VRAM after load; orchestrator aborts on CPU-speed rows and keeps the box for in-place repair |
076c78865 |
bench: wait up to UPSTREAM_WAIT (1 h) for the baseline tarball instead of failing |
a013d01eb |
bench: vast-bench.sh --machine-id (A/B on the same physical box) |
c952cfdd3 |
bench: land picks a free branch name when a run of the same machine landed the same day |
ea72de808 |
docs: release intro notes the removed forward-compatibility library |
79a12ee47 |
bench: benchmark grid on rented multi-GPU machines (Vast.ai); images without cuda-compat (#1) |
License: MIT (upstream llama.cpp license preserved). Upstream llama.cpp is where this work belongs; this fork exists to be archived once upstream performs comparably.