engine 3c7e643 for sm_120 (glibc 2.35+)
A build of this fork at 3c7e64326 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
3c7e64326 is fork main after the engine-5 train: PRs #25-#50 since engine-c008fe8. Among them:
- the PQ2_0 tensor-core kernel for decode and MTP verify rows (#39);
- the MTP draft head's K/V-only catch-up (#42) and its draft vocabulary (
--spec-draft-mtp-vocab, #44); - the scheduler's staged host inputs (#47);
- #50, which stops #39's kernel from holding about 1.4 GB of device memory on a 5070 Ti (1.7 GB on a 5080) for a driver syscall stack.
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (pinned by digest sha256:e7cb1151…; glibc 2.35, gcc 11.4, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX PRO 4000 Blackwell whose container image was that digest.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's 13 ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Checked before upload with rig's scripts/e2e-driver-only.sh, in a fresh Ubuntu 22.04.5 container (glibc 2.35) with one RTX 5070 Ti, driver 610.43.02 and only curl: install.sh; rig prepare finds the prebuilt and no toolkit; rig build installs it and NVIDIA's runtime and passes ldd -r; llama-bench on the ProCreations pack gives tg128 80.6 tok/s.
Measurements: on an RTX 5080, 48 held-out greedy requests at 40,960 tokens of context: rig's bonsai-2-27b head on this build, drafting three tokens a round over its draft vocabulary, decodes 190.2 tok/s against 141.9 for rig 0.1.6's head (two tokens a round) on engine-c008fe8, +36.8 % per request (95 % CI +34.0 to +39.7), 48 of 48 faster. Four slots over a 294,912-token pool peak at 14,012 MiB on that card. TORAD.md lists every change with its measurement and off switch.
sha256 01520ce38d31369146e5bff3f3e75cbde1d735606471ed1367de1bc1bd66a300