Skip to content

engine 3c7e643 for sm_120 (glibc 2.35+)

Choose a tag to compare

@marcospaulo marcospaulo released this 25 Sep 18:42

A build of this fork at 3c7e64326 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.

3c7e64326 is fork main after the engine-5 train: PRs #25-#50 since engine-c008fe8. Among them:

  • the PQ2_0 tensor-core kernel for decode and MTP verify rows (#39);
  • the MTP draft head's K/V-only catch-up (#42) and its draft vocabulary (--spec-draft-mtp-vocab, #44);
  • the scheduler's staged host inputs (#47);
  • #50, which stops #39's kernel from holding about 1.4 GB of device memory on a 5070 Ti (1.7 GB on a 5080) for a driver syscall stack.

It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:

curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b

Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.

How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (pinned by digest sha256:e7cb1151…; glibc 2.35, gcc 11.4, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX PRO 4000 Blackwell whose container image was that digest.

  • CMAKE_CUDA_ARCHITECTURES=120a
  • the AVX2 + FMA + F16C CPU baseline, not -march=native
  • FlashAttention on, CUDA graphs on
  • -ffile-prefix-map, so no build path is recorded
  • no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
  • the compiler's libgomp.so.1 bundled (GCC runtime library exception)

The highest symbol versions across the tarball's 13 ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.

Checked before upload with rig's scripts/e2e-driver-only.sh, in a fresh Ubuntu 22.04.5 container (glibc 2.35) with one RTX 5070 Ti, driver 610.43.02 and only curl: install.sh; rig prepare finds the prebuilt and no toolkit; rig build installs it and NVIDIA's runtime and passes ldd -r; llama-bench on the ProCreations pack gives tg128 80.6 tok/s.

Measurements: on an RTX 5080, 48 held-out greedy requests at 40,960 tokens of context: rig's bonsai-2-27b head on this build, drafting three tokens a round over its draft vocabulary, decodes 190.2 tok/s against 141.9 for rig 0.1.6's head (two tokens a round) on engine-c008fe8, +36.8 % per request (95 % CI +34.0 to +39.7), 48 of 48 faster. Four slots over a 294,912-token pool peak at 14,012 MiB on that card. TORAD.md lists every change with its measurement and off switch.

sha256 01520ce38d31369146e5bff3f3e75cbde1d735606471ed1367de1bc1bd66a300