Skip to content

Releases: neurall/llama.cpp

b11399 glm-speedup

Choose a tag to compare

@neurall neurall released this 28 Sep 10:56

Added

  • LLAMA_GRAPH_REUSE_WHY=1 (GLM-5-next): logs which graph input changed when decode has to rebuild the compute graph instead of reusing it.

Changed

  • MoE expert cache: experts that aren't in VRAM now have id -1 in the device table, and the CUDA MoE mat-vec kernels skip negative ids (they write zeros) instead of computing a dummy slot. Quantized expert weights only; LLAMA_MOE_CACHE_NEG_IDS=0 restores the dummy slot.

Fixed

  • GLM-5-next (GLM-5.3-Flash) decode rebuilt the compute graph every other token. The sparse indexer's key pool completes a pooled key every second token, which changed the graph's shape, so graph reuse (and its CUDA graphs) failed on half the tokens. The pool input now keeps a fixed shape (padded to at least one entry; the padding writes the current token's pooled slot, which is never read), so decode reuses the graph every token. GLM-5.3-Flash, 2 GPUs: chat decode 19.98 -> 21.41 t/s (+7.2%), perplexity run 17.33 -> 18.41 t/s (+6.2%). LLAMA_KPOOL_PAD=0 restores the old behaviour.

Removed

  • Nothing.

Regression check vs b11397 (GLM-5.3-Flash q4kattn, 2 GPUs, model in RAM)

b11397 this release
chat decode (t/s) 19.98 21.41
perplexity run (t/s) 17.33 18.41
perplexity (deterministic cache) 1.4544 1.4519
cache hits (chat) 69.8% 69.8%

Perplexity is unchanged within the MoE cache's run-to-run spread (1.449-1.465 across runs; which experts run on GPU vs CPU varies).

Linux and Windows CUDA builds. On Windows also extract the cudart zip if you don't have CUDA 13.4 installed.

b11397 preheat

Choose a tag to compare

@neurall neurall released this 27 Sep 00:12

Added

  • Prefill preheat. While a long prompt is processed, the experts it uses are copied to the GPU anyway (for the offloaded matmuls); the expert cache now keeps the prompt's most-used ones from those copies, device-to-device, with no extra PCIe traffic, replacing cached experts the prompt used less. Decode after the prompt starts with the prompt's experts in VRAM. GLM-5.3-Flash, 12k-token code prompt then 256 tokens (1 GPU): decode 13.5 -> 15.3 t/s (+12.7%), cache hits 43% -> 55%; prompt processing -1%; 32-token answers and chat unchanged (+1.8%, +0.6%). LLAMA_MOE_CACHE_ADOPT (share of a layer's slots per prompt batch, default 1.0; 0 = off).
  • LLAMA_MOE_CACHE_DROP=1 (opt-in, experimental): for models bigger than RAM with mmap'd weights, drop the RAM pages of experts held in VRAM and read evicted ones back ahead. MiMo-V2.6 on 1 GPU measured -9% (a small, churning cache drops and re-reads too often), so it is off by default.

Changed

  • Nothing.

Fixed

  • Nothing.

Removed

  • Nothing.

Regression check vs b11391 (GLM-5.3-Flash, 1 GPU, model in RAM)

b11391 this release
12k prompt then 256 tokens: decode / hits 13.5 / 43.1% 15.3 / 54.5%
12k prompt then 32 tokens: decode 13.0 13.3
12k prompt processing 269.5 266
chat decode 15.6 15.7

Linux and Windows CUDA builds. On Windows also extract the cudart zip if you don't have CUDA 13.4 installed.

b11391 persistent-hot

Choose a tag to compare

@neurall neurall released this 26 Sep 22:16

Added

  • Persistent hot experts. When the server stops, the expert cache saves how often each expert was used (lifetime activation counts, ~50 KB per model, ~/.cache/llama.cpp/). The next start preloads each layer's most-used experts into VRAM before the first request instead of warming up over the first answers. GLM-5.3-Flash, first request after a restart (2 pairs): decode +4-6% (20.5-20.9 -> 21.8 t/s), cache hits +2-4 points, short-prompt processing ~3x (5.6 -> 17 t/s). Steady-state speed is unchanged. LLAMA_MOE_CACHE_PROFILE=<file> or =0 (off).
  • A GGUF can carry the counts itself (u32 array moe_cache.expert_usage); used when there is no local profile. The model file is never written to.
  • --moe-cache-window N: tokens of recent expert usage the cache scores by (was a fixed 64).
  • LLAMA_MOE_CACHE_STICKY=0.3: keep each layer's most-used 30% of slots from being evicted (off by default: +1.6% measured, within noise).

Changed

  • Experts used by GPU-offloaded prompt batches (-ub 2048) are counted in the usage statistics again (they were missing, so long prompts didn't shape the cache or the profile).

Fixed

  • Nothing.

Removed

  • Nothing.

Regression check vs b11367

Cache scheduling is unchanged; the only runtime difference is the preload at startup and the usage counting. GLM-5.3-Flash chat without a profile (LLAMA_MOE_CACHE_PROFILE=0): 20.5-20.9 t/s, same as b11367 (20.4-20.5).

Tried and not shipped: uploading a long prompt's hot experts right after it (hit rate 67% -> 72%, but the extra uploads compete with decode: 32-token answers -6% to -20%). Next: copy them GPU-to-GPU while the prompt is processed.

Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.

b11367 mtp-fixes

Choose a tag to compare

@neurall neurall released this 26 Sep 17:41

Fixed

  • MTP draft-depth tuner picked the wrong depth by chance. Early in a reply (the model's reasoning) depths measure almost tied, so noise decided; on Qwen3.8-Flash-Next it sometimes settled on depth 1 (~51 t/s) instead of 2 (~60 t/s). The tuner now keeps its current depth (at first the model-size guess) unless another depth is at least 5% faster. Qwen3.8-Flash-Next + MTP chat: 60.4 t/s (2.18x stock), with pinned weights (prompt processing keeps the pinned +45%).

Changed

  • Built-in MTP layers are no longer skipped on models much bigger than VRAM. The ">2x VRAM: don't load a draft" rule now applies only to a separate -md draft (GLM-5.3-Flash's costs the expert cache more than it gains). Built-in MTP, e.g. MiMo-V2.6's 3 dense MTP layers, starts at depth 0 and the tuner decides; on MiMo-V2.6-Flash-RL here it keeps depth 0 (drafting was not 5% faster), so no measurable gain on this box.
  • tools/moe-bench/perf.py: a discarded run before every measured run (model switches and a previous pinned load caused 5-12% swings on first runs); leftover servers are killed by process name.

Added

  • tools/moe-bench/perf.db: every benchmark run behind the README numbers.
  • tools/moe-bench/glm_splice_mtp.py / gguf_remote.py: how the GLM-5.3-Flash MTP head on Hugging Face was made (the MTP tensors are pulled from unsloth's GGUF with HTTP range requests, ~4.3 GiB instead of the whole model).

Removed

  • Nothing.

Regression check vs b11341

Only the MTP depth tuner and defaults changed (no kernels or cache code):

b11341 this release
Qwen chat + MTP head (-md, default settings) ~51 (tuner could pick depth 1) 60.4 (depth 2; 4 more runs: 55.8-62.3, all depth 2)
MiMo chat, --spec-type draft-mtp MTP skipped loaded, tuner keeps depth 0 (no gain)

GLM, MiMo and 12k MTP numbers are being re-measured with the fixed tuner and will be updated in the README.

Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.

b11341 pinning

Choose a tag to compare

@neurall neurall released this 26 Sep 16:18

Added

  • Pinned weights in llama-server: when a MoE model is bigger than VRAM but fits in the RAM available at startup, the server loads its weights into pinned (page-locked) RAM instead of memory-mapping the file. The GPUs then read experts by direct DMA: 12k-token prompt processing GLM-5.3-Flash 179 -> 261 t/s (+46%, now 1.20x stock), Qwen3.8-Flash-Next 372 -> 435 t/s (+17%); decode is unchanged. Costs ~60 s more at startup (GLM: ~100 s vs 38 s) and keeps the model's RAM locked while the server runs, so it pays back after a few long prompts. llama-cli and the other tools keep mmap.
  • --load-mode pin pins in any tool; --load-mode mmap keeps mmap in the server.
  • If pinning leaves too little RAM (>2 GiB swapped while loading or <512 MiB left), the server reloads with mmap and logs why.
  • tools/moe-bench/: the frozen 12k/128k prompts, run.py and perf.py behind the README numbers, so anyone can reproduce them.

Changed

  • Pinned memory is registered with cudaHostRegister instead of cudaMallocHost (about 2x faster to pin), and the loader drops the file's page cache behind it, so the file and its pinned copy never both sit in RAM.
  • The expert-cache upload log says whether it uploads from pinned or pageable memory.

Fixed

  • Nothing.

Removed

  • Nothing.

Regression check vs b11327 (same inputs, model in RAM)

b11327 (mmap) this release
GLM tetris decode 29.5 29.7
GLM chat decode (avg of 2) 20.4 20.5
GLM 12k prompt / decode 180 / 15.6 261 / 16.0 (pinned)
Qwen 12k prompt / decode 372 / 38.4 435 / 39.0 (pinned)

The pinned 12k numbers were measured on test builds with the same pinning code; this release was checked to pin (0 GiB swapped, 12.4 GiB RAM left with GLM). Remaining pinned + MTP numbers will be added to the README as they are measured.

Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.

b11327 qwen3next-mtp

Choose a tag to compare

@neurall neurall released this 26 Sep 13:54

Added

  • Qwen3.8-Flash-Next MTP (PR #28243 by @danielhanchen), working together with the expert cache: 57.0 t/s with MTP vs 47.1 without (1.21x) on a 1500-token chat reply (2x RTX 3090, model already in RAM).
    llama-server -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf -md mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf --spec-type draft-mtp
    
    The MTP head is the small MTP/ file from unsloth/Qwen3.8-Flash-Next-GGUF.
  • GLM-5.3-Flash MTP (PR #27917 by timkhronos), with GLM's MTP head loadable as a separate 4.3 GiB file: neuralll/GLM-5.3-Flash-MTP-GGUF, for any glm5-next GLM-5.3-Flash GGUF, no second copy of the model.
  • Automatic draft depth from measured speed. Each drafted token must be verified, and experts that miss the VRAM cache run on the CPU at full cost per drafted token, so the best depth depends on how much of the model fits in VRAM. The fork starts from a guess (model size vs free VRAM), then measures generation speed and moves the depth up or down. LLAMA_SPEC_DEPTH=N pins it for benchmarks.
  • Draft skipped where it would be slower, with a warning: GLM-5.3-Flash at 2.3x VRAM on this box decodes 20.2 t/s without MTP vs 17.6 with it. --spec-draft-n-max N forces it on. With 3+ GPUs GLM should gain like Qwen.

Changed

  • Default max draft depth: 2 when the model is bigger than free VRAM (the recurrent-state rollback buffers of hybrid models are sized by it and take cache VRAM), 5 when it fits; was 3. --spec-draft-n-max still sets it.
  • The expert cache starts on the first real request instead of at context creation, so it sizes itself after a draft model or mmproj is loaded.
  • Only the main context builds, steps and frees the expert cache; a draft context never touches it.
  • Draft contexts use at most -ub 512 (drafts are a few tokens; 2048 only took cache VRAM).

Fixed

  • Loading a draft model (-md) with the expert cache failed with "unable to allocate CUDA buffer".

Removed

  • Nothing.

Regression check vs b11297 (same inputs, model in RAM)

b11297 this release
GLM tetris decode 28.5 29.6
GLM chat decode (avg of 2) 20.4 20.4
GLM 12k prompt / decode 180 / 15.0 179 / 15.2
Qwen chat decode, no MTP (avg of 2-3) 47.2 47.1
Qwen 12k prompt / decode (avg of 2) 341 / 39.4 372 / 38.4
Qwen chat decode, with MTP (automatic) - (no MTP support) 57.0

Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.

b11297 auto-defaults

Choose a tag to compare

@neurall neurall released this 26 Sep 11:42

Added

  • Automatic defaults that differ from stock llama.cpp, only when a MoE model's weights are larger than the free VRAM of all GPUs (other models and every setting you pass behave as in stock). llama-server -m model.gguf is enough; the chosen values are printed at startup (-lv 5 prints all).
setting stock default fork default why
expert placement autofit: whole layers on GPU, rest on CPU all experts in RAM (--cpu-moe) the VRAM is used for the expert cache instead of fixed layers
expert cache off fills each GPU's free VRAM (--moe-expert-cache -1) holds the experts tokens actually use, so most expert work runs on the GPUs
CPU weight repacking on off (-nr) repacked experts can't be copied to the GPU cache
-ub (tokens per prompt step) 512 2048 if the largest GPU has 20+ GiB free, 1024 at 10+ GiB, else 512 every expert upload serves more prompt tokens: ~2x faster long prompts; costs ~0.7 GiB cache VRAM (~1% decode) on GLM-5.3-Flash
-c (context) model maximum, fitted to VRAM 32768 a bigger KV cache would take VRAM from the expert cache; -c 65536 if you need more
-t (threads) all physical cores cores minus one per GPU a free core per GPU keeps kernel launches and cache uploads fast; measured faster
  • Prompt processing goes to the GPU in the fastest PCIe slot, on Linux and Windows: the fork measures each GPU's upload bandwidth at startup.
  • LLAMA_MOE_CACHE_DETERMINISTIC=1 for reproducible benchmarks.
  • Lookahead expert prefetch from PR #28414 by @leshchukandrej (--prefetch-experts-slots N, off by default).

Changed

  • Decode is faster: the cache is published every step (+16% decode on long prompts), the VRAM safety margin dropped from 1 GB to 384 MiB per GPU, and small batches (up to 31 tokens: short prompts, speculative verification) go through the cache.
  • Layers keep bus order (moving them to the fastest GPU cost ~7% decode on this box; see the README research notes).
  • Latest upstream llama.cpp merged.

Fixed

  • Qwen3.8-Flash-Next segfault on the first request with the cache.
  • MiMo-V2.6 startup assert with the cache; MiMo no longer needs -fitt.

Removed

  • Nothing.

Decode: 1.5x to 2.5x stock on MoE models bigger than VRAM

2x RTX 3090 (one AM4 CPU PCIe 4.0 x16, one X570 chipset x4 slot) + Ryzen 7 3700X + 125 GB DDR4, stock llama.cpp -> this release:

model short prompt, 1500-token chat reply 12k-token code prompt, decode after it
GLM-5.3-Flash 3.0-bit Q4_K attn (106 GB) 13.8 -> 21.5 t/s (1.56x)* 12.3 -> 15.7 t/s (1.28x)*
MiMo-V2.6-Flash-RL IQ3_XXS (132 GB) 4.0 -> 10.1 t/s (2.54x)** 4.2 -> 8.3 t/s (1.98x)**
Qwen3.8-Flash-Next UD-IQ4_XS (88 GB) 27.7 -> 46.4 t/s (1.68x)* 25.3 -> 38.7 t/s (1.53x)*

* Model already in RAM (OS page cache), as on a server after its first request. The first run after switching to another large model is slower, once, while the file is read from disk. ** MiMo (132 GB) can't fully stay cached in 125 GB RAM.

Models that fit in VRAM (Qwen3.8-27B, OLMoE) are unaffected: same speed, identical output.

Prompt processing is still up to ~22% slower than stock (stock keeps whole layers in VRAM); using both GPUs' PCIe links for it is the next milestone.

Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.

b11214 expert-cache

Choose a tag to compare

@neurall neurall released this 24 Sep 14:12

Portable Linux CUDA build (GLM-5.3-Flash + VRAM-filling MoE expert cache fork). See README for benchmarks.

Important for prompt processing (prefill) speed: put the GPU in your fastest PCIe slot first. Prefill uploads the experts to the first GPU, and GPUs are numbered by PCI bus order, not slot speed, so the first one is often a card in a slow chipset x4 slot. This binary does not reorder automatically; set it yourself, e.g. CUDA_VISIBLE_DEVICES=1,0 llama-server ... (Linux) or set CUDA_VISIBLE_DEVICES=1,0 (Windows) when GPU 1 is the x16 card. Check with nvidia-smi --query-gpu=index,pcie.link.width.current --format=csv under load. On 2x RTX 3090 (x4 + x16) this took a 12k-token prompt from 45.9 to 84.1 t/s, and to 227.7 t/s adding -ub 2048 -b 2048.