Skip to content

engine 4104c47 for sm_120 (glibc 2.35+)

Choose a tag to compare

@marcospaulo marcospaulo released this 27 Sep 10:14

A build of this fork at 4104c47 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.

4104c47 is the train engine-10 (#70), 4 commits past engine-02512a3. The ternary PQ2_0 matmuls stream closer to the card's DRAM rate, and the engine says where a decode token's time goes:

  • Each PQ2_0 launch prefetches the next one's head into L2 (16ad036). A launch's blocks take every gridDim.x-th 16-row tile, so what a launch reads first is its matrices' heads. Once a block's last weight box has landed, the block asks L2 (cp.async.bulk.prefetch.L2.global) for its share of the next PQ2_0 launch's head, the gate's too for a gated pair, while the small kernels between the two launches run. The size is planned per launch from the card's profile: 2 us of its DRAM rate, a quarter of its L2 at most (GGML_CUDA_PQ2_PREFETCH_US, 0 turns it off).
  • The device's profile at init (SMs, clocks, registers and shared memory per SM, L2 and its persisting share, the DRAM rate from the memory clock and bus width) and, with GGML_CUDA_NVTX=1, an NVTX range around every node that launches work, named with the bytes it reads and writes, for Nsight Systems (1c83edd).
  • llama-bench -cts, the recurrent state's cache type, as an axis like -ctk/-ctv (b2eb433); -ctk, -ctv and -cts take f32 (4104c47).

It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:

curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b

Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.

How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.

  • CMAKE_CUDA_ARCHITECTURES=120a
  • the AVX2 + FMA + F16C CPU baseline, not -march=native
  • FlashAttention on, CUDA graphs on
  • -ffile-prefix-map, so no build path is recorded
  • no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
  • the compiler's libgomp.so.1 bundled (GCC runtime library exception)

Measurements:

  • Ternary Bonsai 2 27B, llama-bench tg64 as rig serves it (q4_0 K/V, f16 state, CUDA graphs on), 4 rounds × 5 repetitions a cell with the order rotated each round and a discarded warm-up, medians, the prefetch at 2 us against the same build's parent (b2eb433):
    • RTX 5080: +2.42 % at depth 16,384, +2.70 % at depth 0;
    • RTX 5070 Ti: +1.56 % and +1.10 %;
    • RTX 5090: +4.84 % and +4.51 % (152.93 → 160.34 and 157.57 → 164.68 tok/s), against engine-02512a3's libggml-cuda under this build's binaries.
      The greedy text is byte-identical with the prefetch on and off. test-backend-ops passes its 296 PQ2_0 and 24 MUL_MAT_GROUP cases at 0, 8 and 64 us.
  • This build against engine-02512a3's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with ldd before any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities: bit-exact, plain and with the MTP draft (the same 1,536 tokens, log-probabilities, top-5 and draft counts); without n_probs, the same tokens and drafting; a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.01 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041). The legs' decode rates, one run each in the gate's order: plain 90.9-91.8 → 92.8-94.0 tok/s, drafted 143.3-168.3 → 144.8-170.9.
  • scripts/e2e-driver-only.sh with this tarball on an RTX 5070 Ti (Ubuntu 22.04, driver only): install, ldd -r, decode pass.

sha256 20d9f06e33a19a81b9ee47af844b165ae4970c26b81a305d0c86a0ad50af5741