engine 4283c36 for sm_120 (glibc 2.35+)
A build of this fork at 4283c36bc for NVIDIA Blackwell GeForce and workstation cards (sm_120: RTX 5080, 5070 Ti, 5090, RTX PRO 6000). It holds llama-server, llama-bench, llama-kv-mean-center and their libraries, with RUNPATH $ORIGIN. It runs on glibc 2.35 or newer: Ubuntu 22.04+, Debian 12+, Fedora 36+.
4283c36bc is the train engine-10 (#70), 27 commits past engine-c1518d4. Among them:
- the recurrent state runs in f16 and bf16 (
-cts f16,-cts bf16; 1fd214c, 48ebd21). A state row was zeroed withGGML_OP_SCALE, which CUDA and the CPU run on f32 only, so either type aborted on its first decode. The server writes each sequence's state back in its format after every call, and on Ternary Bonsai 2 27B a q8_0 state's error accumulates over a long decode; an f16 state measures as f32 within noise; - the tokens a draft verification emits carry their probabilities (a8d2290);
- a JSON schema's
{}below its root accepts any value, not only an object, so a free-form tool argument parses under the lazy tool grammar, with GBNF and llguidance (737eba9, c206d74); - a slot requested by id loads its prompt from the cache when empty (07e92c7), and a resumed task takes the checkpoint it resumed from as its own (ab8f91a);
- saving and restoring a transposed V computes its offsets in 64 bits (9123de6); a restore that throws logs why (df4f8e5, f4e0972);
- behind their own off switches: a decode's flash attention leaves out the mask scans that cannot skip anything at one sequence per stream (6c0dc9b,
GGML_CUDA_FATTN_MASK_PREFIX_LEGACY), the attention rotations are set once in the cache's buffer (b0dd41c,LLAMA_ATTN_ROT_INPUT_LEGACY), and a chain whose top-k picks first takes each row's top k on the backend (6c0e372,LLAMA_TOP_K_PREFILTER_LEGACY); every*_LEGACYswitch reads 0, false, no and off as off (4384bbb, 198e77e); - a sampler leaving the backend drops the graph built with it (83e3c09): graph reuse compared samplers by address, and the prefilter's chain, made at each launch with its k in the graph, could reuse the previous request's graph and draw from its k;
- draft-MTP options given without
--spec-type draft-mtpare refused at load instead of dropped, and-bswith--pull-layerswarns (4283c36); - the scheduler stages inputs only where its async set runs in the compute's queue (67b853d), and the NUMA bind counts the draft model's threads (d5f5eb7).
It is what rig installs when engine/engine.toml pins this release. rig checks the file against the sha256 below. rig also fetches NVIDIA's CUDA 13.3 runtime from NVIDIA's own server (cuda_cudart 13.3.29, libcublas 13.5.1.27, each pinned by sha256) and unpacks it beside the binaries. A machine needs only an NVIDIA driver that supports CUDA 13; no toolkit or compiler:
curl -fsSL https://github.com/torad-labs/rig/releases/latest/download/install.sh | sh
rig up bonsai-2-27b
Using it without rig: unpack the tarball. Put libcudart.so.13, libcublas.so.13 and libcublasLt.so.13 from a CUDA 13.x install (or NVIDIA's redistributable archives) beside the binaries or on LD_LIBRARY_PATH.
How it was built: rig build --portable in nvidia/cuda:13.3.0-devel-ubuntu22.04 (pulled by tag on the rented box; glibc 2.35, gcc 11.4.0 and CUDA 13.3.33, the contents of rig's pinned digest sha256:e7cb1151…), with the two steps of rig's scripts/prebuilt/Dockerfile on top (CMake 3.31.6 by sha256, ninja). It ran on a rented RTX 5090.
CMAKE_CUDA_ARCHITECTURES=120a- the AVX2 + FMA + F16C CPU baseline, not
-march=native - FlashAttention on, CUDA graphs on
-ffile-prefix-map, so no build path is recorded- no OpenSSL: llama-server's HTTPS downloads and TLS serving are off; serve local files on loopback
- the compiler's
libgomp.so.1bundled (GCC runtime library exception)
The highest symbol versions across the tarball's ELF files are GLIBC_2.34, GLIBCXX_3.4.30 and CXXABI_1.3.13.
Checked before upload with rig's scripts/e2e-driver-only.sh: a fresh ubuntu:22.04 container (Ubuntu 22.04.5, glibc 2.35) given one RTX 5080 and driver 610.43.02 through CDI, nothing else from NVIDIA; install.sh; rig prepare finds the prebuilt and no toolkit; rig build installs it and NVIDIA's runtime and passes ldd -r, every CUDA library resolved from the build directory; llama-bench on the ProCreations pack gives tg128 103.7 tok/s.
Measurements: engine-48ebd21, this build less its last two commits, against engine-c1518d4's on a rented RTX 5090, each with NVIDIA's pinned runtime (checked with ldd before any leg), one slot at 262,144, on a 245,752-token prompt's four questions, greedy, 384 tokens each with top-5 log-probabilities:
- with its three new switches at legacy, bit-exact against engine-c1518d4, plain and with the MTP draft: the same 1,536 tokens, log-probabilities and draft counts;
- as served, against engine-c1518d4: three first differences, each at a near-tie (the largest 0.111 nats), |dlogprob| over the 961 tokens before them p99 0.078 and max 0.181;
- its MTP draft (three tokens a round) against its own plain decode: four first differences, each at a near-tie (the largest 0.244 nats), |dlogprob| p99 0.096 and max 0.174 over 348 tokens; every drafted token carries its probabilities (engine-c1518d4: 1,532 of 1,536 without);
- without
n_probs, where the top-k prefilter runs, the same tokens and drafting as with it, plain and drafted; - a 245K conversation swapped out of its slot and back at 4 × 294,912 returns its first token in 0.93 s with the answer of the conversation never swapped out, token for token (|dlogprob| p99 0.016);
- a tool whose parameter is
{"type": "object", "additionalProperties": {}}: the served call carries the model's own flat schema under the lazy tool grammar, where engine-c1518d4's grammar changes it.
The f16 state, measured on the source pack over 8,192 decoded tokens after 131,072, three tokens a call against f16 K/V and an f32 state: the served K/V with a q8_0 state 0.002488 mean KLD, with an f16 state 0.001724 (the difference 0.000764 ± 0.000070, batch-means); an f16 state costs 6.3 % a draft round against q8_0 on an RTX 5080 at 245K (95 % CI 6.2 to 6.4).
This build against engine-48ebd21, on an RTX 5080 (driver 610.43.02), each with NVIDIA's pinned runtime (checked with ldd before any leg), the head's argv at one slot at 262,144, the same prompt's four questions, greedy, 384 tokens each: plain and drafted, each with top-5 log-probabilities and without n_probs: the same 1,536 tokens on every leg, the same log-probabilities and top-5 at every token of the two legs that ask for them, and the same draft counts on both drafted legs. The measurements above hold for this build. With the prefilter at legacy (LLAMA_TOP_K_PREFILTER_LEGACY=1) and no n_probs, the same tokens and drafting as with it, plain and drafted, and the same speed within one run's noise (the four questions' mean: plain 62.8 tok/s at legacy against 63.2 with the prefilter, drafted 113.0 against 113.3): at one slot the decode waits on the card, and the prefilter's gain is the server thread's CPU time.
sha256 4848503632f4416b3a3a4aaa1139a5202d8040065b6c3220e9430d2464e3d883