engine c1518d4 for sm_120 (glibc 2.35+)
A build of this fork at c1518d4cd for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
c1518d4cd is fork main after the trains engine-6 (#56), engine-7 (#62), engine-8 (#64) and engine-9 (#69): PRs #51-#69 since engine-3c7e643. Among them:
- flash attention's stream-k splits only the KV steps each Q tile sees, at decode and prefill (#63). With
--kv-unifiedit split the pool's whole span, so a sequence in the far part of a large pool left most of its blocks without work; - the KQ mask built from each sequence's cells as bits, 64 cells a word (#65);
- rms_norm, its weight multiply and the signed FWHT in one kernel at decode (#51, #58);
- a restore into scattered cells copies a run of cells at a time (#68), and a resumed prompt copies no checkpoint twice (#67);
- the attention gate and the Gated DeltaNet output read in place: 64 copy kernels a step fewer on Ternary Bonsai 2 27B (engine-9);
- on a host with more than one NUMA node, llama-server runs on the node of the model's GPUs (#55).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (pinned by digest sha256:e7cb1151…; glibc 2.35, gcc 11.4, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX PRO 4000 Blackwell whose container image was that digest.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's 13 ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Checked before upload with rig's scripts/e2e-driver-only.sh, its container's script run on a rented box whose image is Ubuntu 22.04.5 (glibc 2.35), with one RTX 5090 and driver 610.57.04 (the host's ssh bootstrap adds git and python3 to the image; git was removed before the script's check, and no step of the install runs python3): install.sh; rig prepare finds the prebuilt and no toolkit; rig build installs it and NVIDIA's runtime and passes ldd -r; llama-bench on the ProCreations pack gives tg128 161.1 tok/s.
Measurements: an RTX 5090 serving rig's bonsai-2-27b head as rig 0.1.7 sizes it there: eight slots over one unified pool of 786,432 cells, three draft tokens a round over the draft vocabulary, sampled at temperature 1.0. Three conversations of 178,000 tokens fill the pool's first cells, then a 245,760-token prompt lands in its last. engine-3c7e643 prefilled that prompt in 364.7 s and decoded it at 22.31 ms a draft round (99.6 tok/s). This build took 116.8 s at 13.32 ms (165.2 tok/s). With the prompt in the pool's first cells, the two read 117.3 s at 14.75 ms and 115.2 s at 13.15 ms. The three conversations before it prefilled in 72.8, 132.4 and 192.3 s on engine-3c7e643, and in 71.7, 71.9 and 72.1 s here. On an RTX 5080 this build was checked against engine-3c7e643's, one slot at 262,144, on the 245,760-token prompt's four questions greedy with top-5 log-probabilities. Every first difference is at a near-tie (0.011 to 0.137 nats), and plain decode there reads 58.6 against 47.3-48.0 tok/s. With the draft, the text differs from plain only at near-ties. A conversation swapped out to host memory and back returns its first token in 3.2 s with the same answer. TORAD.md lists every change with its measurement and off switch.
sha256 c41ac41cf83f1b4a20ee0101afddefaf3d98ccd149f9c52e6a13cecd0fe86854