engine 48ebd21 for sm_120 (glibc 2.35+)
A build of this fork at 48ebd2167 for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
48ebd2167 is the train engine-10, 25 commits past engine-c1518d4. Among them:
- the recurrent state runs in f16 and bf16 (
-cts f16,-cts bf16; 1fd214c, 48ebd21). A state row was zeroed withGGML_OP_SCALE, which CUDA and the CPU run on f32 only, so either type aborted on its first decode. The server writes each sequence's state back in its format after every call, and on Ternary Bonsai 2 27B a q8_0 state's error accumulates over a long decode; an f16 state measures as f32 within noise; - the tokens a draft verification emits carry their probabilities (a8d2290);
- a JSON schema's
{}below its root accepts any value, not only an object, so a free-form tool argument parses under the lazy tool grammar, with GBNF and llguidance (737eba9, c206d74); - a slot requested by id loads its prompt from the cache when empty (07e92c7), and a resumed task takes the checkpoint it resumed from as its own (ab8f91a);
- saving and restoring a transposed V computes its offsets in 64 bits (9123de6); a restore that throws logs why (df4f8e5, f4e0972);
- behind their own off switches: a decode's flash attention leaves out the mask scans that cannot skip anything at one sequence per stream (6c0dc9b,
GGML_CUDA_FATTN_MASK_PREFIX_LEGACY), the attention rotations are set once in the cache's buffer (b0dd41c,LLAMA_ATTN_ROT_INPUT_LEGACY), and a chain whose top-k picks first takes each row's top k on the backend (6c0e372,LLAMA_TOP_K_PREFILTER_LEGACY); every*_LEGACYswitch reads 0, false, no and off as off (4384bbb, 198e77e); - the scheduler stages inputs only where its async set runs in the compute's queue (67b853d), and the NUMA bind counts the draft model's threads (d5f5eb7).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (pulled by tag on the rented box; glibc 2.35, gcc 11.4.0 and CUDA 13.3.33, the contents of rig's pinned digest sha256:e7cb1151…), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Checked before upload with rig's scripts/e2e-driver-only.sh: a fresh ubuntu:22.04 container (Ubuntu 22.04.5, glibc 2.35) given one RTX 5080 and driver 610.43.02 through CDI, nothing else from NVIDIA; install.sh; rig prepare finds the prebuilt and no toolkit; rig build installs it and NVIDIA's runtime and passes ldd -r, every CUDA library resolved from the build directory; llama-bench on the ProCreations pack gives tg128 104.4 tok/s.
Measurements: this tarball against engine-c1518d4's on a rented RTX 5090, each with NVIDIA's pinned runtime (checked with ldd before any leg), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities:
- with its three new switches at legacy, bit-exact against engine-c1518d4, plain and with the MTP draft: the same 1,536 tokens, log-probabilities and draft counts;
- as served, against engine-c1518d4: three first differences, each at a near-tie (the largest 0.111 nats), |dlogprob| over the 961 tokens before them p99 0.078 and max 0.181;
- its MTP draft (three tokens a round) against its own plain decode: four first differences, each at a near-tie (the largest 0.244 nats), |dlogprob| p99 0.096 and max 0.174 over 348 tokens; every drafted token carries its probabilities (engine-c1518d4: 1,532 of 1,536 without);
- without
n_probs, where the top-k prefilter runs, the same tokens and drafting as with it, plain and drafted; - a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 0.93 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.016);
- a tool whose parameter is
{"type": "object", "additionalProperties": {}}: the served call carries the model's own flat schema under the lazy tool grammar, where engine-c1518d4's grammar changes it.
The f16 state, measured on the source pack over 8,192 decoded tokens after 131,072, three tokens a call against f16 K/V and an f32 state: the served K/V with a q8_0 state 0.002488 mean KLD, with an f16 state 0.001724 (the difference 0.000764 ± 0.000070, batch-means); an f16 state costs 6.3 % a draft round against q8_0 on an RTX 5080 at 245K (95 % CI 6.2 to 6.4).
sha256 378ba695e244eeea7d1942ecfa0a60706c6cf034b6fdc692f173c3dff9e0f1cc