engine c008fe8 for sm_120 (glibc 2.35+)
A build of this fork at c008fe8cf for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
c008fe8cf has the same source tree as 13998300f (tree 6d0ff0bc), the commit of engine-1399830-2. This fork's history was rewritten on 2026-09-24 to remove assistant-attribution lines from commit messages, which gave every commit a new SHA and changed no file. rig pins this commit from its next release on; the older releases stay as published, because rig v0.1.0 and v0.1.1 pin them by sha256.
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig's scripts/build-prebuilt.sh, which runs rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (pinned by digest; glibc 2.35, gcc 11.4, CUDA 13.3.33).
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's 13 ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Checked before upload with rig's scripts/e2e-driver-only.sh, in a fresh Ubuntu 22.04.5 container (glibc 2.35) with one RTX 5070 Ti, driver 610.43.02 and only curl: install.sh; rig prepare finds the prebuilt and no toolkit; rig build installs it and NVIDIA's runtime and passes ldd -r; llama-bench on the ProCreations pack gives tg128 73.3 tok/s.
Measurements: this commit is the served engine. docs/torad/bonsai-27b has the numbers, and TORAD.md lists every change with its off switch.
sha256 12aeb3b24497dec4b1a8536585c66ee585ebd82ed13d663a1b9d82e9b812a8d5