Skip to content

engine 87a3596 for sm_120 (glibc 2.35+)

Choose a tag to compare

@marcospaulo marcospaulo released this 27 Sep 03:20

A build of this fork at 87a359611 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.

87a359611 is the train engine-10 (#70), 3 commits past engine-15a502b:

  • a graph per decode shape, each with its own scheduler: a speculative verify that alternates shapes (an MTP draft's catch-up and draft graphs, an n-gram draft's rows) reuses its graphs instead of rebuilding them every round (9ddd463). LLAMA_GRAPH_REUSE_DISABLE keeps the old path;
  • the Gated DeltaNet kernel reads and writes an f16 recurrent cache itself, as it does q8_0, instead of a gather to f32 and a copy back per layer (9760e96). GGML_CUDA_GDN_F16_CACHE_LEGACY=1 keeps the copies;
  • at load each context logs what its graph slots can add to its compute buffers, graph slot buffer size <= X MiB, 3 slots (87a3596).

It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:

curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b

Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.

How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.

  • CMAKE_CUDA_ARCHITECTURES=120a
  • the AVX2 + FMA + F16C CPU baseline, not -march=native
  • FlashAttention on, CUDA graphs on
  • -ffile-prefix-map, so no build path is recorded
  • no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
  • the compiler's libgomp.so.1 bundled (GCC runtime library exception)

The highest symbol versions across the tarball's ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.

Measurements: this build against engine-15a502b's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with ldd before any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities:

  • bit-exact, plain and with the MTP draft: the same 1,536 tokens, log-probabilities, top-5 and draft counts;
  • without n_probs, where the top-k prefilter runs, the same tokens and drafting, plain and drafted;
  • a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.00 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041).

Served speed on an RTX 5080 (4 × 294,912, f16 state, MTP draft n_max 3, the draft vocabulary), six agentic requests of 94K–121K prompt tokens, greedy, 512 tokens each, one binary through the switches: the f16 cache fusion 185.77 → 202.09 tok/s (same-text geomean +8.81 %, se 0.30 %; every request +7.3 % to +9.9 %; the same text and draft counters on all six), the round 18.11 → 16.65 ms; the graph slots +0.65 % (se 0.17 %) with the MTP draft alone and +1.55 % (se 0.35 %) with n-gram drafts beside it. The graph slots hold up to 3 graphs of up to 32 rows per context: +234 MiB at peak on the 5080 under four concurrent requests, inside the 241 MiB the logged bound gives.

sha256 91264ffd104b730e54913639e567d4bfa7e7bf10660dff0eae1242ec079336ec