Skip to content

engine 4b61c54 for sm_120 (glibc 2.35+)

Choose a tag to compare

@marcospaulo marcospaulo released this 27 Sep 17:56

A build of this fork at 4b61c54 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.

4b61c54 is on the train engine-10 branch, 6 commits past engine-4104c47. The raw q4_0/q8_0 flash attention stops losing time to its own shared memory at depth:

  • Conflict-free shared memory in the raw flash attention (01f4f0f). At 65,536 cells on an RTX 5080 (Ternary Bonsai 2 27B's attention: head 256, 4 KV heads at GQA 6, a bit-packed mask), 3.43M of the kernel's 7.36M shared-memory wavefronts were bank conflicts: the V tile's dequant stored 16 bytes from 8 threads on 2 rows (4-way), and the K scales' float tile loaded 32 rows of one block a warp (4-way). The dequant's store phase now spans 8 rows (store conflicts 3.2M → 0.09M). K·Q reads each block's f16 scale from the raw rows, dropping the float tile, its pass and its barrier (load conflicts 0.23M → 0.03M; 26,272 bytes a block, so a decode token runs 3 blocks an SM where it ran 2). The stream-k fixup loads the next 8 blocks' partials before folding (6.9 → 5.9 us), padding columns write no fixup partial, and the next raw K tile loads beside V where 2 blocks share an SM. GGML_CUDA_FATTN_Q4_0_LEGACY=1 / GGML_CUDA_FATTN_Q8_0_LEGACY=1 restore the stock kernels.
  • llama-bench -rs, the recurrent-state snapshots a draft's rollback keeps (7537f40). A drafting server runs n_rs_seq = its draft's n_max, so every verify of a hybrid model writes n_max + 1 snapshots of each recurrent layer's state; llama-bench wrote one. On an RTX 5080, pp4 at depth 16,384 in a 512 ubatch: -rs 0 386.15, -rs 3 383.76 tok/s.
  • A tensor split's ratio failure names its node (6b98c5d): the node, op, dims, axes and segments, and its source chain, where a bare assert printed nothing. Fatal path only.

It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:

curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b

Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.

How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (glibc 2.35, gcc 11.4.0, CUDA 13.3.33), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.

  • CMAKE_CUDA_ARCHITECTURES=120a
  • the AVX2 + FMA + F16C CPU baseline, not -march=native
  • FlashAttention on, CUDA graphs on
  • -ffile-prefix-map, so no build path is recorded
  • no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
  • the compiler's libgomp.so.1 bundled (GCC runtime library exception)

Measurements:

  • The served attention in isolation (test-backend-ops perf, q4_0 K/V read raw with a bit-packed mask, the cache's layout), RTX 5080, 6 rounds against the parent's library, a token / a 4-row verify:
    • 16,384 cells: 30.95 → 25.78 us (-16.7 %) / 33.02 → 30.80 (-6.7 %);
    • 65,536: 109.44 → 105.44 (-3.7 %) / 114.79 → 110.66 (-3.6 %);
    • 131,072: 195.62 → 188.14 (-3.8 %) / 204.07 → 198.19 (-2.9 %);
    • 245,760: 341.41 → 335.17 (-1.8 %) / 355.46 → 345.67 (-2.8 %).
      FLASH_ATTN_EXT passes 3,220/3,220; greedy 64 tokens after a 7K-token prompt are byte-identical to the parent's.
  • This build against engine-4104c47's on a rented RTX 5090 (driver 615.71.09), each with NVIDIA's pinned runtime (checked with ldd before any leg), Ternary Bonsai 2 27B as rig serves it (q4_0 K/V with its mean-center, f16 recurrent state), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities. A decode token's tile running 3 blocks an SM moves its stream-k split (340 → 510 ways on a 5090): the same arithmetic summed in another order. So the pairs are held to the bars engine-10's gate held its own reordering to (every first difference at a tie, |dlogprob| p99 ≤ 0.15 and max ≤ 0.5), not to bit-exactness: plain, 4 first differences, each at a tie, and |dlogprob| p99 0.056, max 0.090 over the 326 agreeing tokens before them; with the MTP draft, one first difference, at a tie, p99 0.032, max 0.112 over 1,221; without n_probs, the same tokens and drafting (bit-exact); a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 1.03 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.027, max 0.041). The legs' decode rates, one run each in the gate's order: plain 91.9-93.2 → 93.0-94.5 tok/s, the 245K question itself 91.9 → 93.7.
  • scripts/e2e-driver-only.sh with this tarball on an RTX 5070 Ti (Ubuntu 22.04, driver only): install, ldd -r, decode pass.

sha256 ebba0ca80438b494d336fd059e7637d43bb817b18a7e360d01c6340549628b0d