engine 15a502b for sm_120 (glibc 2.35+)
A build of this fork at 15a502bc5 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
15a502bc5 is the train engine-10 (#70), 3 commits past engine-a29b719:
- the MMA flash attention reads a q8_0 K/V cache raw, as it reads q4_0: K and V straight from the cache, once per GQA group, K·Q on int8 tensor cores (15a502b).
GGML_CUDA_FATTN_Q8_0_LEGACY=1restores the stock kernels; - ngram-mod finds an n-gram only in the slot it filled (bc59cc7);
seq_rmrefuses a rollback past the snapshots the last batch wrote (6510894).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Measurements: this build against engine-a29b719's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with ldd before any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities:
- bit-exact, plain and with the MTP draft: the same 1,536 tokens, log-probabilities, top-5 and draft counts;
- without
n_probs, where the top-k prefilter runs, the same tokens and drafting, plain and drafted; - a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 0.97 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041).
The same prompt against an f16 K/V cache: q8_0 read raw agrees on 831 of 1,536 tokens before each question's first difference, every difference at a tie (≤ 0.3 nats), |dlogprob| p99 0.035; the stock q8_0 kernels 760, p99 0.034; q4_0 as served 366, one difference beyond a tie (0.313 nats), p99 0.126.
llama-bench on the same card, the stock kernels against the raw path through the switch, one binary: q8_0 decode at a 245,760-token depth 44.95 → 83.74 tok/s (1.86×), prefill pp512 at 65,536 2,214.63 → 2,800.09 tok/s; q4_0 unchanged (103.22 → 103.29 tok/s decode at 245,760).
sha256 f517ff414f0164b26def5e6bd76b62942eeef48d629b9c705ac37158b5e4494e