A build of this fork at 32e695e for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
32e695e is on the train engine-10 branch, 5 commits past engine-4b61c54. It fixes that release's RTX 5090 regression at depth and speeds up the multi-slot flash attention on every card:
- The live-tile split's blocks aligned to a Q tile's output tiles (32e695e). With more than one slot and
--kv-unified, engine-4b61c54 split a decode token's attention over every block the SMs hold. At 3 blocks an SM that is 510 blocks on an RTX 5090 (127.5 per KV head) and 210 on an RTX 5070 Ti (52.5), so each head's blocks read the cache half a block apart from its neighbour's. A cell's K and V rows hold the heads side by side (144 bytes each) and the L2 fetches 64-byte units, so neighbouring heads share a unit, and it was fetched twice: the cache was read 4/3 times (ncu on a 5070 Ti at 131,072 cells: 202-207 MB of DRAM reads for 151 MB of K and V). The split now rounds down to a multiple of the output tiles (508 and 208 blocks): 152-153 MB.GGML_CUDA_FATTN_LIVE_BLOCKS_ALIGN_LEGACY=1restores every block. - One copy of the mma flash attention's tile code (818965d). The tile code was compiled once per stream-k role, 184 KB of SASS in three copies. Only the 4 blocks that end a tile ran their copy, cold in every cache after the K/V stream, and they ended 8-9 us after the rest. The roles are now arguments of one 71 KB copy (libggml-cuda 79.1 → 73.2 MB). The same arithmetic and stores; no switch.
- The live steps scanned once a graph (6aa1467). The live-tile split scanned the mask before each of the model's 16 attention layers, 3.2-4.5 us each, though every layer reads the same mask. The graph's first scan is kept, keyed by the mask and the split.
GGML_CUDA_FATTN_LIVE_SCAN_EACH_LEGACY=1scans every layer. - Pre-sampling probabilities ride the top-k prefilter (f1e45c2). A request for
n_probs(and every OpenAIlogprobsrequest) turned the prefilter off, so every token copied the row's 248,320 logits to the host for a CPU softmax.llama_sampler_init_row_probs()takes the row's softmax on the backend.LLAMA_TOP_K_PREFILTER_LEGACY=1keeps every logit on the CPU.
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
Measurements:
- A decode token's attention on an RTX 5090 (
test-backend-ops perf, q4_0 K/V read raw with a bit-packed mask, the cache's layout, 6 rounds, one binary over each release's libraries), us at 16,384 / 65,536 / 131,072 / 245,760 cells:- engine-4104c47: 26.89 / 51.46 / 123.69 / 204.66;
- engine-4b61c54: 28.84 / 44.57 / 145.50 / 255.48;
- this commit (its code built on the box): 19.60 / 37.84 / 115.72 / 196.44.
A 4-row verify: 28.83 / 56.30 / 131.97 / 216.32 → 21.62 / 47.73 / 122.24 / 207.04.
- The served decode on the RTX 5090 at 245K tokens with 4 slots (
-c 294912 -np 4 --kv-unified), the 245K prompt's four questions, 256 greedy tokens, 3 legs each, this build with and without its alignment: plain 98.75 → 107.92 tok/s (+9.3 %), with the MTP draft 205.43 → 208.90 (+1.7 %). At that shape with top-5 log-probabilities, the two splits' texts part at ties: plain 4 first differences, each at a tie, |dlogprob| p99 0.095, max 0.128; drafted 4, p99 0.070, max 0.078. - With
n_probs5 on an RTX 5080 at -c 32768: plain 95.65 → 105.14 tok/s, drafted 166.20 → 192.51, the same tokens and top-5 ids. - This build against engine-4b61c54's on a rented RTX 5090, each with NVIDIA's pinned runtime, Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities. The aligned split sums the same arithmetic in another order, so the pairs are held to engine-10's G1 bars (every first difference at a tie, |dlogprob| p99 ≤ 0.15 and max ≤ 0.5), not bit-exactness: plain, 4 first differences, each at a tie, p99 0.080, max 0.144 over the 323 agreeing tokens before them; with the MTP draft, 4, each at a tie, p99 0.089, max 0.227 over 601; without
n_probs, the same tokens and drafting (bit-exact); a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.08 s with the answer of the conversation never swapped out, token for token. Decode on the 245K question with top-5 probabilities asked: 94.3 → 107.5 tok/s plain. scripts/e2e-driver-only.shwith this tarball on an RTX 5070 Ti (Ubuntu 22.04, driver only): install,ldd -r, decode pass (tg128 87.33 tok/s).
sha256 9df73c20b9b70695b2c39fb592aeae7ee1ab6c4af5db134ae16b321e990b1c21