Skip to content

engine 1399830 — sm_120 (Blackwell) build

Choose a tag to compare

@marcospaulo marcospaulo released this 24 Sep 17:16
· 91 commits to 13998300f70388c1fbc865b411a975a8c1f84cfe since this release

A build of this fork at 13998300f for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN.

It is what rig installs when engine/engine.toml pins this commit. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13 — no toolkit, no compiler:

curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b

Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.

How it was built: rig build --portable with rig v0.1.0, and checked by its scripts/e2e-driver-only.sh (Ubuntu 24.04, driver only: install, ldd -r, decode) before upload:

  • CUDA 13.3 with CMAKE_CUDA_ARCHITECTURES=120a
  • the AVX2 + FMA + F16C CPU baseline, not -march=native
  • FlashAttention on, CUDA graphs on
  • -ffile-prefix-map, so no build path is recorded
  • no OpenSSL (llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback)
  • the compiler's libgomp.so.1 bundled (GCC runtime library exception); glibc 2.38 or newer (Ubuntu 24.04+, Debian 13+, Fedora 39+)

Measurements: this commit is the served engine. docs/torad/bonsai-27b has the numbers, and TORAD.md lists every change with its off switch.

sha256 fdaaa51bc4769d20b418c148e8c55f349e70bffcfb884b7e9f32f66cb92741f1