Skip to content

engine 02512a3 for sm_120 (glibc 2.35+)

Choose a tag to compare

@marcospaulo marcospaulo released this 27 Sep 07:43

A build of this fork at 02512a3 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.

02512a3 is the train engine-10 (#70), 5 commits past engine-b9e1b78. A decode step reuses the graph it built the step before instead of rebuilding it wherever the inputs allow:

  • GLM-5.3's pooled indexer input (llm_graph_input_kpool) had no reuse check, and an input without one refuses, so every GLM-5.3 decode step rebuilt its whole graph (eec6d1d). Its check compares every shape the builder sizes from the step's geometry, with the builder's own pool count (7a17ff5);
  • a recurrent input compares rs_z only when a graph baked it: an f32 state's in-graph zeroing is the only place it enters a graph, and a context with no recurrent layers (GLM-5.3's NextN draft context) never reads it (02512a3). Before, such a context held a second copy of each draft graph, one per rs_z;
  • with LLAMA_GRAPH_INPUT_DEBUG=2 and -lv, a refused reuse names its input and the check that refused (bf76114, cc9e415, 02512a3), and a Release build logs the CUDA-graph refusals it used to compile out: a MUL_MAT_ID that needs a stream sync, with its node, and a failed executable update (7a17ff5).

It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:

curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b

Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.

How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.

  • CMAKE_CUDA_ARCHITECTURES=120a
  • the AVX2 + FMA + F16C CPU baseline, not -march=native
  • FlashAttention on, CUDA graphs on
  • -ffile-prefix-map, so no build path is recorded
  • no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
  • the compiler's libgomp.so.1 bundled (GCC runtime library exception)

Measurements:

  • GLM-5.3 IQ3_XXS on 2 × RTX PRO 6000 (bf76114, which holds the kpool check), the same fixed probes at temperature 0: decode 101.5 and 108.5 tok/s against 87 at b9e1b78; with graph reuse turned off (LLAMA_GRAPH_REUSE_DISABLE=1) 75.2, and the same tokens and top-5 log-probabilities at all 150 positions as with it on.
  • Ternary Bonsai 2 27B drafted on a rented RTX 5090, 512 tokens at temperature 0: 285.9 and 286.5 tok/s; 214.6 with graph reuse off and 287.2 with CUDA graphs off, the same text in every run. Its reuse was already whole: the 512 tokens replay 9 graphs.
  • This build against engine-b9e1b78's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with ldd before any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities: bit-exact, plain and with the MTP draft (the same 1,536 tokens, log-probabilities, top-5 and draft counts); without n_probs, the same tokens and drafting; a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.01 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041).
  • scripts/e2e-driver-only.sh with this tarball on an RTX 5070 Ti (Ubuntu 22.04, driver only): install, ldd -r, decode pass.

The rs_z change is measured on bonsai (unchanged there: its f32 conv state bakes rs_z) and read on GLM-5.3's log, where every rs_z refusal was followed by a second graph for the same draft shape; its GLM-5.3 run is still to come.

sha256 7966c1fa9ecb2d50d1210d85df572a1e80f61e050b020df0ad0fd7611e37e470