engine 1399830 for sm_120 (glibc 2.35+)
A build of this fork at 13998300f for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN.
This is the second build of 13998300f. It is the same source as engine-1399830, built inside Ubuntu 22.04, so it runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+. engine-1399830 needs glibc 2.38; rig v0.1.0 and v0.1.1 pin it by sha256, so it stays as published.
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig's scripts/build-prebuilt.sh, which runs rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (pinned by digest; glibc 2.35, gcc 11.4, CUDA 13.3.33).
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across all 16 ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Checked before upload with rig's scripts/e2e-driver-only.sh, in fresh containers with one RTX 5070 Ti, driver 610.43.02 and only curl:
- Ubuntu 22.04.5 (glibc 2.35): install.sh;
rig preparefinds the prebuilt and no toolkit;rig buildinstalls it and NVIDIA's runtime and passesldd -r; llama-bench on the ProCreations pack gives pp512 1,800 and tg128 73.2 tok/s. - Ubuntu 24.04.5 (glibc 2.39): install and
ldd -rpass.
Measurements: this commit is the served engine. docs/torad/bonsai-27b has the numbers, and TORAD.md lists every change with its off switch.
sha256 4a6d6ed6838b8e8a054bf56ac7a00ce74d9c999efd0e4550ad38f7ab99eb676e