Releases: torad-labs/llama.cpp
Release list
engine b9e1b78 for sm_120 (glibc 2.35+)
A build of this fork at b9e1b78ec for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
b9e1b78ec is the train engine-10 (#70), 2 commits past engine-87a3596:
- at most 8 CUDA graphs per context, the least recently used evicted (d0f8bae): the graphs one per shape were bounded only by a 10 s eviction, so the memory they held followed the traffic.
GGML_CUDA_GRAPH_MAXsets the cap (0: none); the cap is logged at load, and the most held and the largest instance as they grow; test-recurrent-rollback-rounding(b9e1b78): a recurrent state rolled back holds to the rounding of its batch shapes, in the state's own type (f32, f16, q8_0), and one token early or late fails it.
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Measurements: this build against engine-87a3596's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with ldd before any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities:
- bit-exact, plain and with the MTP draft: the same 1,536 tokens, log-probabilities, top-5 and draft counts;
- without
n_probs, where the top-k prefilter runs, the same tokens and drafting, plain and drafted; - a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.03 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041).
The cap's cost, on an RTX 5080 at 4 × 294,912 under four concurrent requests: peak 15,180 MiB capped against 15,228 uncapped; one request at a time, the same text at 214.6 tok/s against 216.3, the five requests after the first within 0.3 %. scripts/e2e-driver-only.sh with this tarball on an RTX 5070 Ti (Ubuntu 22.04, driver only): install, ldd -r, decode pass.
sha256 8e1de8e13e68d7d66997baa9c4edba950533de5e7e06031d4b32e0ad4ef37af2
engine 87a3596 for sm_120 (glibc 2.35+)
A build of this fork at 87a359611 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
87a359611 is the train engine-10 (#70), 3 commits past engine-15a502b:
- a graph per decode shape, each with its own scheduler: a speculative verify that alternates shapes (an MTP draft's catch-up and draft graphs, an n-gram draft's rows) reuses its graphs instead of rebuilding them every round (9ddd463).
LLAMA_GRAPH_REUSE_DISABLEkeeps the old path; - the Gated DeltaNet kernel reads and writes an f16 recurrent cache itself, as it does q8_0, instead of a gather to f32 and a copy back per layer (9760e96).
GGML_CUDA_GDN_F16_CACHE_LEGACY=1keeps the copies; - at load each context logs what its graph slots can add to its compute buffers,
graph slot buffer size <= X MiB, 3 slots(87a3596).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Measurements: this build against engine-15a502b's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with ldd before any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities:
- bit-exact, plain and with the MTP draft: the same 1,536 tokens, log-probabilities, top-5 and draft counts;
- without
n_probs, where the top-k prefilter runs, the same tokens and drafting, plain and drafted; - a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.00 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041).
Served speed on an RTX 5080 (4 × 294,912, f16 state, MTP draft n_max 3, the draft vocabulary), six agentic requests of 94K–121K prompt tokens, greedy, 512 tokens each, one binary through the switches: the f16 cache fusion 185.77 → 202.09 tok/s (same-text geomean +8.81 %, se 0.30 %; every request +7.3 % to +9.9 %; the same text and draft counters on all six), the round 18.11 → 16.65 ms; the graph slots +0.65 % (se 0.17 %) with the MTP draft alone and +1.55 % (se 0.35 %) with n-gram drafts beside it. The graph slots hold up to 3 graphs of up to 32 rows per context: +234 MiB at peak on the 5080 under four concurrent requests, inside the 241 MiB the logged bound gives.
sha256 91264ffd104b730e54913639e567d4bfa7e7bf10660dff0eae1242ec079336ec
engine 4b61c54 for sm_120 (glibc 2.35+)
A build of this fork at 4b61c54 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
4b61c54 is on the train engine-10 branch, 6 commits past engine-4104c47. The raw q4_0/q8_0 flash attention stops losing time to its own shared memory at depth:
- Conflict-free shared memory in the raw flash attention (01f4f0f). At 65,536 cells on an RTX 5080 (Ternary Bonsai 2 27B's attention: head 256, 4 KV heads at GQA 6, a bit-packed mask), 3.43M of the kernel's 7.36M shared-memory wavefronts were bank conflicts: the V tile's dequant stored 16 bytes from 8 threads on 2 rows (4-way), and the K scales' float tile loaded 32 rows of one block a warp (4-way). The dequant's store phase now spans 8 rows (store conflicts 3.2M → 0.09M). K·Q reads each block's f16 scale from the raw rows, dropping the float tile, its pass and its barrier (load conflicts 0.23M → 0.03M; 26,272 bytes a block, so a decode token runs 3 blocks an SM where it ran 2). The stream-k fixup loads the next 8 blocks' partials before folding (6.9 → 5.9 us), padding columns write no fixup partial, and the next raw K tile loads beside V where 2 blocks share an SM.
GGML_CUDA_FATTN_Q4_0_LEGACY=1/GGML_CUDA_FATTN_Q8_0_LEGACY=1restore the stock kernels. - llama-bench
-rs, the recurrent-state snapshots a draft's rollback keeps (7537f40). A drafting server runsn_rs_seq= its draft's n_max, so every verify of a hybrid model writes n_max + 1 snapshots of each recurrent layer's state; llama-bench wrote one. On an RTX 5080, pp4 at depth 16,384 in a 512 ubatch:-rs 0386.15,-rs 3383.76 tok/s. - A tensor split's ratio failure names its node (6b98c5d): the node, op, dims, axes and segments, and its source chain, where a bare assert printed nothing. Fatal path only.
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
Measurements:
- The served attention in isolation (
test-backend-ops perf, q4_0 K/V read raw with a bit-packed mask, the cache's layout), RTX 5080, 6 rounds against the parent's library, a token / a 4-row verify:- 16,384 cells: 30.95 → 25.78 us (-16.7 %) / 33.02 → 30.80 (-6.7 %);
- 65,536: 109.44 → 105.44 (-3.7 %) / 114.79 → 110.66 (-3.6 %);
- 131,072: 195.62 → 188.14 (-3.8 %) / 204.07 → 198.19 (-2.9 %);
- 245,760: 341.41 → 335.17 (-1.8 %) / 355.46 → 345.67 (-2.8 %).
FLASH_ATTN_EXTpasses 3,220/3,220; greedy 64 tokens after a 7K-token prompt are byte-identical to the parent's.
- This build against engine-4104c47's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with
lddbefore any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities. A decode token's tile running 3 blocks an SM moves its stream-k split (340 → 510 ways on a 5090): the same arithmetic summed in another order. So the pairs are held to the bars engine-10's gate held its own reordering to (every first difference at a tie, |dlogprob| p99 ≤ 0.15 and max ≤ 0.5), not to bit-exactness: plain, 4 first differences, each at a tie, and |dlogprob| p99 0.056, max 0.090 over the 326 agreeing tokens before them; with the MTP draft, one first difference, at a tie, p99 0.032, max 0.112 over 1,221; withoutn_probs, the same tokens and drafting (bit-exact); a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.03 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041). The legs' decode rates, one run each in the gate's order: plain 91.9-93.2 → 93.0-94.5 tok/s, the 245K question itself 91.9 → 93.7. scripts/e2e-driver-only.shwith this tarball on an RTX 5070 Ti (Ubuntu 22.04, driver only): install,ldd -r, decode pass.
sha256 ebba0ca80438b494d336fd059e7637d43bb817b18a7e360d01c6340549628b0d
engine 4104c47 for sm_120 (glibc 2.35+)
A build of this fork at 4104c47 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
4104c47 is the train engine-10 (#70), 4 commits past engine-02512a3. The ternary PQ2_0 matmuls stream closer to the card's DRAM rate, and the engine says where a decode token's time goes:
- Each PQ2_0 launch prefetches the next one's head into L2 (16ad036). A launch's blocks take every gridDim.x-th 16-row tile, so what a launch reads first is its matrices' heads. Once a block's last weight box has landed, the block asks L2 (
cp.async.bulk.prefetch.L2.global) for its share of the next PQ2_0 launch's head, the gate's too for a gated pair, while the small kernels between the two launches run. The size is planned per launch from the card's profile: 2 us of its DRAM rate, a quarter of its L2 at most (GGML_CUDA_PQ2_PREFETCH_US, 0 turns it off). - The device's profile at init (SMs, clocks, registers and shared memory per SM, L2 and its persisting share, the DRAM rate from the memory clock and bus width) and, with
GGML_CUDA_NVTX=1, an NVTX range around every node that launches work, named with the bytes it reads and writes, for Nsight Systems (1c83edd). - llama-bench
-cts, the recurrent state's cache type, as an axis like-ctk/-ctv(b2eb433);-ctk,-ctvand-ctstakef32(4104c47).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
Measurements:
- Ternary Bonsai 2 27B,
llama-benchtg64 as rig serves it (q4_0 K/V, f16 state, CUDA graphs on), 4 rounds × 5 repetitions a cell with the order rotated each round and a discarded warm-up, medians, the prefetch at 2 us against the same build's parent (b2eb433):- RTX 5080: +2.42 % at depth 16,384, +2.70 % at depth 0;
- RTX 5070 Ti: +1.56 % and +1.10 %;
- RTX 5090: +4.84 % and +4.51 % (152.93 → 160.34 and 157.57 → 164.68 tok/s), against engine-02512a3's
libggml-cudaunder this build's binaries.
The greedy text is byte-identical with the prefetch on and off. test-backend-ops passes its 296 PQ2_0 and 24 MUL_MAT_GROUP cases at 0, 8 and 64 us.
- This build against engine-02512a3's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with
lddbefore any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities: bit-exact, plain and with the MTP draft (the same 1,536 tokens, log-probabilities, top-5 and draft counts); withoutn_probs, the same tokens and drafting; a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.01 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041). The legs' decode rates, one run each in the gate's order: plain 90.9-91.8 → 92.8-94.0 tok/s, drafted 143.3-168.3 → 144.8-170.9. scripts/e2e-driver-only.shwith this tarball on an RTX 5070 Ti (Ubuntu 22.04, driver only): install,ldd -r, decode pass.
sha256 20d9f06e33a19a81b9ee47af844b165ae4970c26b81a305d0c86a0ad50af5741
engine 32e695e for sm_120 (glibc 2.35+)
A build of this fork at 32e695e for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
32e695e is on the train engine-10 branch, 5 commits past engine-4b61c54. It fixes that release's RTX 5090 regression at depth and speeds up the multi-slot flash attention on every card:
- The live-tile split's blocks aligned to a Q tile's output tiles (32e695e). With more than one slot and
--kv-unified, engine-4b61c54 split a decode token's attention over every block the SMs hold. At 3 blocks an SM that is 510 blocks on an RTX 5090 (127.5 per KV head) and 210 on an RTX 5070 Ti (52.5), so each head's blocks read the cache half a block apart from its neighbour's. A cell's K and V rows hold the heads side by side (144 bytes each) and the L2 fetches 64-byte units, so neighbouring heads share a unit, and it was fetched twice: the cache was read 4/3 times (ncu on a 5070 Ti at 131,072 cells: 202-207 MB of DRAM reads for 151 MB of K and V). The split now rounds down to a multiple of the output tiles (508 and 208 blocks): 152-153 MB.GGML_CUDA_FATTN_LIVE_BLOCKS_ALIGN_LEGACY=1restores every block. - One copy of the mma flash attention's tile code (818965d). The tile code was compiled once per stream-k role, 184 KB of SASS in three copies. Only the 4 blocks that end a tile ran their copy, cold in every cache after the K/V stream, and they ended 8-9 us after the rest. The roles are now arguments of one 71 KB copy (libggml-cuda 79.1 → 73.2 MB). The same arithmetic and stores; no switch.
- The live steps scanned once a graph (6aa1467). The live-tile split scanned the mask before each of the model's 16 attention layers, 3.2-4.5 us each, though every layer reads the same mask. The graph's first scan is kept, keyed by the mask and the split.
GGML_CUDA_FATTN_LIVE_SCAN_EACH_LEGACY=1scans every layer. - Pre-sampling probabilities ride the top-k prefilter (f1e45c2). A request for
n_probs(and every OpenAIlogprobsrequest) turned the prefilter off, so every token copied the row's 248,320 logits to the host for a CPU softmax.llama_sampler_init_row_probs()takes the row's softmax on the backend.LLAMA_TOP_K_PREFILTER_LEGACY=1keeps every logit on the CPU.
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
Measurements:
- A decode token's attention on an RTX 5090 (
test-backend-ops perf, q4_0 K/V read raw with a bit-packed mask, the cache's layout, 6 rounds, one binary over each release's libraries), us at 16,384 / 65,536 / 131,072 / 245,760 cells:- engine-4104c47: 26.89 / 51.46 / 123.69 / 204.66;
- engine-4b61c54: 28.84 / 44.57 / 145.50 / 255.48;
- this commit (its code built on the box): 19.60 / 37.84 / 115.72 / 196.44.
A 4-row verify: 28.83 / 56.30 / 131.97 / 216.32 → 21.62 / 47.73 / 122.24 / 207.04.
- The served decode on the RTX 5090 at 245K tokens with 4 slots (
-c 294912 -np 4 --kv-unified), the 245K prompt's four questions, 256 greedy tokens, 3 legs each, this build with and without its alignment: plain 98.75 → 107.92 tok/s (+9.3 %), with the MTP draft 205.43 → 208.90 (+1.7 %). At that shape with top-5 log-probabilities, the two splits' texts part at ties: plain 4 first differences, each at a tie, |dlogprob| p99 0.095, max 0.128; drafted 4, p99 0.070, max 0.078. - With
n_probs5 on an RTX 5080 at -c 32768: plain 95.65 → 105.14 tok/s, drafted 166.20 → 192.51, the same tokens and top-5 ids. - This build against engine-4b61c54's on a rented RTX 5090, each with NVIDIA's pinned runtime, Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities. The aligned split sums the same arithmetic in another order, so the pairs are held to engine-10's G1 bars (every first difference at a tie, |dlogprob| p99 ≤ 0.15 and max ≤ 0.5), not bit-exactness: plain, 4 first differences, each at a tie, p99 0.080, max 0.144 over the 323 agreeing tokens before them; with the MTP draft, 4, each at a tie, p99 0.089, max 0.227 over 601; without
n_probs, the same tokens and drafting (bit-exact); a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.08 s with the answer of the conversation never swapped out, token for token. Decode on the 245K question with top-5 probabilities asked: 94.3 → 107.5 tok/s plain. scripts/e2e-driver-only.shwith this tarball on an RTX 5070 Ti (Ubuntu 22.04, driver only): install,ldd -r, decode pass (tg128 87.33 tok/s).
sha256 9df73c20b9b70695b2c39fb592aeae7ee1ab6c4af5db134ae16b321e990b1c21
engine 15a502b for sm_120 (glibc 2.35+)
A build of this fork at 15a502bc5 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
15a502bc5 is the train engine-10 (#70), 3 commits past engine-a29b719:
- the MMA flash attention reads a q8_0 K/V cache raw, as it reads q4_0: K and V straight from the cache, once per GQA group, K·Q on int8 tensor cores (15a502b).
GGML_CUDA_FATTN_Q8_0_LEGACY=1restores the stock kernels; - ngram-mod finds an n-gram only in the slot it filled (bc59cc7);
seq_rmrefuses a rollback past the snapshots the last batch wrote (6510894).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Measurements: this build against engine-a29b719's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with ldd before any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities:
- bit-exact, plain and with the MTP draft: the same 1,536 tokens, log-probabilities, top-5 and draft counts;
- without
n_probs, where the top-k prefilter runs, the same tokens and drafting, plain and drafted; - a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 0.97 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041).
The same prompt against an f16 K/V cache: q8_0 read raw agrees on 831 of 1,536 tokens before each question's first difference, every difference at a tie (≤ 0.3 nats), |dlogprob| p99 0.035; the stock q8_0 kernels 760, p99 0.034; q4_0 as served 366, one difference beyond a tie (0.313 nats), p99 0.126.
llama-bench on the same card, the stock kernels against the raw path through the switch, one binary: q8_0 decode at a 245,760-token depth 44.95 → 83.74 tok/s (1.86×), prefill pp512 at 65,536 2,214.63 → 2,800.09 tok/s; q4_0 unchanged (103.22 → 103.29 tok/s decode at 245,760).
sha256 f517ff414f0164b26def5e6bd76b62942eeef48d629b9c705ac37158b5e4494e
engine 02512a3 for sm_120 (glibc 2.35+)
A build of this fork at 02512a3 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
02512a3 is the train engine-10 (#70), 5 commits past engine-b9e1b78. A decode step reuses the graph it built the step before instead of rebuilding it wherever the inputs allow:
- GLM-5.3's pooled indexer input (
llm_graph_input_kpool) had no reuse check, and an input without one refuses, so every GLM-5.3 decode step rebuilt its whole graph (eec6d1d). Its check compares every shape the builder sizes from the step's geometry, with the builder's own pool count (7a17ff5); - a recurrent input compares
rs_zonly when a graph baked it: an f32 state's in-graph zeroing is the only place it enters a graph, and a context with no recurrent layers (GLM-5.3's NextN draft context) never reads it (02512a3). Before, such a context held a second copy of each draft graph, one perrs_z; - with
LLAMA_GRAPH_INPUT_DEBUG=2and-lv, a refused reuse names its input and the check that refused (bf76114, cc9e415, 02512a3), and a Release build logs the CUDA-graph refusals it used to compile out: a MUL_MAT_ID that needs a stream sync, with its node, and a failed executable update (7a17ff5).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
Measurements:
- GLM-5.3 IQ3_XXS on 2 × RTX PRO 6000 (bf76114, which holds the kpool check), the same fixed probes at temperature 0: decode 101.5 and 108.5 tok/s against 87 at b9e1b78; with graph reuse turned off (
LLAMA_GRAPH_REUSE_DISABLE=1) 75.2, and the same tokens and top-5 log-probabilities at all 150 positions as with it on. - Ternary Bonsai 2 27B drafted on a rented RTX 5090, 512 tokens at temperature 0: 285.9 and 286.5 tok/s; 214.6 with graph reuse off and 287.2 with CUDA graphs off, the same text in every run. Its reuse was already whole: the 512 tokens replay 9 graphs.
- This build against engine-b9e1b78's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with
lddbefore any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities: bit-exact, plain and with the MTP draft (the same 1,536 tokens, log-probabilities, top-5 and draft counts); withoutn_probs, the same tokens and drafting; a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.01 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041). scripts/e2e-driver-only.shwith this tarball on an RTX 5070 Ti (Ubuntu 22.04, driver only): install,ldd -r, decode pass.
The rs_z change is measured on bonsai (unchanged there: its f32 conv state bakes rs_z) and read on GLM-5.3's log, where every rs_z refusal was followed by a second graph for the same draft shape; its GLM-5.3 run is still to come.
sha256 7966c1fa9ecb2d50d1210d85df572a1e80f61e050b020df0ad0fd7611e37e470
engine c1518d4 for sm_120 (glibc 2.35+)
A build of this fork at c1518d4cd for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
c1518d4cd is fork main after the trains engine-6 (#56), engine-7 (#62), engine-8 (#64) and engine-9 (#69): PRs #51-#69 since engine-3c7e643. Among them:
- flash attention's stream-k splits only the KV steps each Q tile sees, at decode and prefill (#63). With
--kv-unifiedit split the pool's whole span, so a sequence in the far part of a large pool left most of its blocks without work; - the KQ mask built from each sequence's cells as bits, 64 cells a word (#65);
- rms_norm, its weight multiply and the signed FWHT in one kernel at decode (#51, #58);
- a restore into scattered cells copies a run of cells at a time (#68), and a resumed prompt copies no checkpoint twice (#67);
- the attention gate and the Gated DeltaNet output read in place: 64 copy kernels a step fewer on Ternary Bonsai 2 27B (engine-9);
- on a host with more than one NUMA node, llama-server runs on the node of the model's GPUs (#55).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (pinned by digest sha256:e7cb1151…; glibc 2.35, gcc 11.4, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX PRO 4000 Blackwell whose container image was that digest.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's 13 ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Checked before upload with rig's scripts/e2e-driver-only.sh, its container's script run on a rented box whose image is Ubuntu 22.04.5 (glibc 2.35), with one RTX 5090 and driver 610.57.04 (the host's ssh bootstrap adds git and python3 to the image; git was removed before the script's check, and no step of the install runs python3): install.sh; rig prepare finds the prebuilt and no toolkit; rig build installs it and NVIDIA's runtime and passes ldd -r; llama-bench on the ProCreations pack gives tg128 161.1 tok/s.
Measurements: an RTX 5090 serving rig's bonsai-2-27b head as rig 0.1.7 sizes it there: eight slots over one unified pool of 786,432 cells, three draft tokens a round over the draft vocabulary, sampled at temperature 1.0. Three conversations of 178,000 tokens fill the pool's first cells, then a 245,760-token prompt lands in its last. engine-3c7e643 prefilled that prompt in 364.7 s and decoded it at 22.31 ms a draft round (99.6 tok/s). This build took 116.8 s at 13.32 ms (165.2 tok/s). With the prompt in the pool's first cells, the two read 117.3 s at 14.75 ms and 115.2 s at 13.15 ms. The three conversations before it prefilled in 72.8, 132.4 and 192.3 s on engine-3c7e643, and in 71.7, 71.9 and 72.1 s here. On an RTX 5080 this build was checked against engine-3c7e643's, one slot at 262,144, on the 245,760-token prompt's four questions greedy with top-5 log-probabilities. Every first difference is at a near-tie (0.011 to 0.137 nats), and plain decode there reads 58.6 against 47.3-48.0 tok/s. With the draft, the text differs from plain only at near-ties. A conversation swapped out to host memory and back returns its first token in 3.2 s with the same answer. TORAD.md lists every change with its measurement and off switch.
sha256 c41ac41cf83f1b4a20ee0101afddefaf3d98ccd149f9c52e6a13cecd0fe86854
engine a29b719 for sm_120 (glibc 2.35+)
A build of this fork at a29b71955 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
a29b71955 is the train engine-10 (#70), 4 commits past engine-4283c36:
- GLM-5.3 Flash (
glm5next: Kimi delta attention and DeepSeek sparse attention, one NextN head), ported from ggml-org#27754 at 86ebfef onto this fork's APIs (5531e53, 6037ef1). Its architecture fixture passestest-llama-archs; --spec-rollback Nsets how long a draft a recurrent target rolls back in place (17a4374);- a verify batch writes only the conv-state windows a rollback can read, not all n_rs_seq + 1 of them (in 5531e53; its test, a29b719).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Measurements: this build against engine-4283c36's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with ldd before any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities:
- bit-exact, plain and with the MTP draft: the same 1,536 tokens, log-probabilities, top-5 and draft counts;
- without
n_probs, where the top-k prefilter runs, the same tokens and drafting, plain and drafted; - a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.01 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041).
sha256 6ebe2ed77cbbb92cf0a37381aad12b75226439a63b7404d9633b95e5b234cd42
engine 48ebd21 for sm_120 (glibc 2.35+)
A build of this fork at 48ebd2167 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
48ebd2167 is the train engine-10, 25 commits past engine-c1518d4. Among them:
- the recurrent state runs in f16 and bf16 (
-cts f16,-cts bf16; 1fd214c, 48ebd21). A state row was zeroed withGGML_OP_SCALE, which CUDA and the CPU run on f32 only, so either type aborted on its first decode. The server writes each sequence's state back in its format after every call, and on Ternary Bonsai 2 27B a q8_0 state's error accumulates over a long decode; an f16 state measures as f32 within noise; - the tokens a draft verification emits carry their probabilities (a8d2290);
- a JSON schema's
{}below its root accepts any value, not only an object, so a free-form tool argument parses under the lazy tool grammar, with GBNF and llguidance (737eba9, c206d74); - a slot requested by id loads its prompt from the cache when empty (07e92c7), and a resumed task takes the checkpoint it resumed from as its own (ab8f91a);
- saving and restoring a transposed V computes its offsets in 64 bits (9123de6); a restore that throws logs why (df4f8e5, f4e0972);
- behind their own off switches: a decode's flash attention leaves out the mask scans that cannot skip anything at one sequence per stream (6c0dc9b,
GGML_CUDA_FATTN_MASK_PREFIX_LEGACY), the attention rotations are set once in the cache's buffer (b0dd41c,LLAMA_ATTN_ROT_INPUT_LEGACY), and a chain whose top-k picks first takes each row's top k on the backend (6c0e372,LLAMA_TOP_K_PREFILTER_LEGACY); every*_LEGACYswitch reads 0, false, no and off as off (4384bbb, 198e77e); - the scheduler stages inputs only where its async set runs in the compute's queue (67b853d), and the NUMA bind counts the draft model's threads (d5f5eb7).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (pulled by tag on the rented box; glibc 2.35, gcc 11.4.0 and CUDA 13.3.33, the contents of rig's pinned digest sha256:e7cb1151…), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Checked before upload with rig's scripts/e2e-driver-only.sh: a fresh ubuntu:22.04 container (Ubuntu 22.04.5, glibc 2.35) given one RTX 5080 and driver 610.43.02 through CDI, nothing else from NVIDIA; install.sh; rig prepare finds the prebuilt and no toolkit; rig build installs it and NVIDIA's runtime and passes ldd -r, every CUDA library resolved from the build directory; llama-bench on the ProCreations pack gives tg128 104.4 tok/s.
Measurements: this tarball against engine-c1518d4's on a rented RTX 5090, each with NVIDIA's pinned runtime (checked with ldd before any leg), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities:
- with its three new switches at legacy, bit-exact against engine-c1518d4, plain and with the MTP draft: the same 1,536 tokens, log-probabilities and draft counts;
- as served, against engine-c1518d4: three first differences, each at a near-tie (the largest 0.111 nats), |dlogprob| over the 961 tokens before them p99 0.078 and max 0.181;
- its MTP draft (three tokens a round) against its own plain decode: four first differences, each at a near-tie (the largest 0.244 nats), |dlogprob| p99 0.096 and max 0.174 over 348 tokens; every drafted token carries its probabilities (engine-c1518d4: 1,532 of 1,536 without);
- without
n_probs, where the top-k prefilter runs, the same tokens and drafting as with it, plain and drafted; - a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 0.93 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.016);
- a tool whose parameter is
{"type": "object", "additionalProperties": {}}: the served call carries the model's own flat schema under the lazy tool grammar, where engine-c1518d4's grammar changes it.
The f16 state, measured on the source pack over 8,192 decoded tokens after 131,072, three tokens a call against f16 K/V and an f32 state: the served K/V with a q8_0 state 0.002488 mean KLD, with an f16 state 0.001724 (the difference 0.000764 ± 0.000070, batch-means); an f16 state costs 6.3 % a draft round against q8_0 on an RTX 5080 at 245K (95 % CI 6.2 to 6.4).
sha256 378ba695e244eeea7d1942ecfa0a60706c6cf034b6fdc692f173c3dff9e0f1cc