Releases: neurall/llama.cpp
Release list
b11399 glm-speedup
Added
LLAMA_GRAPH_REUSE_WHY=1(GLM-5-next): logs which graph input changed when decode has to rebuild the compute graph instead of reusing it.
Changed
- MoE expert cache: experts that aren't in VRAM now have id -1 in the device table, and the CUDA MoE mat-vec kernels skip negative ids (they write zeros) instead of computing a dummy slot. Quantized expert weights only;
LLAMA_MOE_CACHE_NEG_IDS=0restores the dummy slot.
Fixed
- GLM-5-next (GLM-5.3-Flash) decode rebuilt the compute graph every other token. The sparse indexer's key pool completes a pooled key every second token, which changed the graph's shape, so graph reuse (and its CUDA graphs) failed on half the tokens. The pool input now keeps a fixed shape (padded to at least one entry; the padding writes the current token's pooled slot, which is never read), so decode reuses the graph every token. GLM-5.3-Flash, 2 GPUs: chat decode 19.98 -> 21.41 t/s (+7.2%), perplexity run 17.33 -> 18.41 t/s (+6.2%).
LLAMA_KPOOL_PAD=0restores the old behaviour.
Removed
- Nothing.
Regression check vs b11397 (GLM-5.3-Flash q4kattn, 2 GPUs, model in RAM)
| b11397 | this release | |
|---|---|---|
| chat decode (t/s) | 19.98 | 21.41 |
| perplexity run (t/s) | 17.33 | 18.41 |
| perplexity (deterministic cache) | 1.4544 | 1.4519 |
| cache hits (chat) | 69.8% | 69.8% |
Perplexity is unchanged within the MoE cache's run-to-run spread (1.449-1.465 across runs; which experts run on GPU vs CPU varies).
Linux and Windows CUDA builds. On Windows also extract the cudart zip if you don't have CUDA 13.4 installed.
b11397 preheat
Added
- Prefill preheat. While a long prompt is processed, the experts it uses are copied to the GPU anyway (for the offloaded matmuls); the expert cache now keeps the prompt's most-used ones from those copies, device-to-device, with no extra PCIe traffic, replacing cached experts the prompt used less. Decode after the prompt starts with the prompt's experts in VRAM. GLM-5.3-Flash, 12k-token code prompt then 256 tokens (1 GPU): decode 13.5 -> 15.3 t/s (+12.7%), cache hits 43% -> 55%; prompt processing -1%; 32-token answers and chat unchanged (+1.8%, +0.6%).
LLAMA_MOE_CACHE_ADOPT(share of a layer's slots per prompt batch, default 1.0; 0 = off). LLAMA_MOE_CACHE_DROP=1(opt-in, experimental): for models bigger than RAM with mmap'd weights, drop the RAM pages of experts held in VRAM and read evicted ones back ahead. MiMo-V2.6 on 1 GPU measured -9% (a small, churning cache drops and re-reads too often), so it is off by default.
Changed
- Nothing.
Fixed
- Nothing.
Removed
- Nothing.
Regression check vs b11391 (GLM-5.3-Flash, 1 GPU, model in RAM)
| b11391 | this release | |
|---|---|---|
| 12k prompt then 256 tokens: decode / hits | 13.5 / 43.1% | 15.3 / 54.5% |
| 12k prompt then 32 tokens: decode | 13.0 | 13.3 |
| 12k prompt processing | 269.5 | 266 |
| chat decode | 15.6 | 15.7 |
Linux and Windows CUDA builds. On Windows also extract the cudart zip if you don't have CUDA 13.4 installed.
b11391 persistent-hot
Added
- Persistent hot experts. When the server stops, the expert cache saves how often each expert was used (lifetime activation counts, ~50 KB per model,
~/.cache/llama.cpp/). The next start preloads each layer's most-used experts into VRAM before the first request instead of warming up over the first answers. GLM-5.3-Flash, first request after a restart (2 pairs): decode +4-6% (20.5-20.9 -> 21.8 t/s), cache hits +2-4 points, short-prompt processing ~3x (5.6 -> 17 t/s). Steady-state speed is unchanged.LLAMA_MOE_CACHE_PROFILE=<file>or=0(off). - A GGUF can carry the counts itself (u32 array
moe_cache.expert_usage); used when there is no local profile. The model file is never written to. --moe-cache-window N: tokens of recent expert usage the cache scores by (was a fixed 64).LLAMA_MOE_CACHE_STICKY=0.3: keep each layer's most-used 30% of slots from being evicted (off by default: +1.6% measured, within noise).
Changed
- Experts used by GPU-offloaded prompt batches (
-ub 2048) are counted in the usage statistics again (they were missing, so long prompts didn't shape the cache or the profile).
Fixed
- Nothing.
Removed
- Nothing.
Regression check vs b11367
Cache scheduling is unchanged; the only runtime difference is the preload at startup and the usage counting. GLM-5.3-Flash chat without a profile (LLAMA_MOE_CACHE_PROFILE=0): 20.5-20.9 t/s, same as b11367 (20.4-20.5).
Tried and not shipped: uploading a long prompt's hot experts right after it (hit rate 67% -> 72%, but the extra uploads compete with decode: 32-token answers -6% to -20%). Next: copy them GPU-to-GPU while the prompt is processed.
Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.
b11367 mtp-fixes
Fixed
- MTP draft-depth tuner picked the wrong depth by chance. Early in a reply (the model's reasoning) depths measure almost tied, so noise decided; on Qwen3.8-Flash-Next it sometimes settled on depth 1 (~51 t/s) instead of 2 (~60 t/s). The tuner now keeps its current depth (at first the model-size guess) unless another depth is at least 5% faster. Qwen3.8-Flash-Next + MTP chat: 60.4 t/s (2.18x stock), with pinned weights (prompt processing keeps the pinned +45%).
Changed
- Built-in MTP layers are no longer skipped on models much bigger than VRAM. The ">2x VRAM: don't load a draft" rule now applies only to a separate
-mddraft (GLM-5.3-Flash's costs the expert cache more than it gains). Built-in MTP, e.g. MiMo-V2.6's 3 dense MTP layers, starts at depth 0 and the tuner decides; on MiMo-V2.6-Flash-RL here it keeps depth 0 (drafting was not 5% faster), so no measurable gain on this box. tools/moe-bench/perf.py: a discarded run before every measured run (model switches and a previous pinned load caused 5-12% swings on first runs); leftover servers are killed by process name.
Added
tools/moe-bench/perf.db: every benchmark run behind the README numbers.tools/moe-bench/glm_splice_mtp.py/gguf_remote.py: how the GLM-5.3-Flash MTP head on Hugging Face was made (the MTP tensors are pulled from unsloth's GGUF with HTTP range requests, ~4.3 GiB instead of the whole model).
Removed
- Nothing.
Regression check vs b11341
Only the MTP depth tuner and defaults changed (no kernels or cache code):
| b11341 | this release | |
|---|---|---|
Qwen chat + MTP head (-md, default settings) |
~51 (tuner could pick depth 1) | 60.4 (depth 2; 4 more runs: 55.8-62.3, all depth 2) |
MiMo chat, --spec-type draft-mtp |
MTP skipped | loaded, tuner keeps depth 0 (no gain) |
GLM, MiMo and 12k MTP numbers are being re-measured with the fixed tuner and will be updated in the README.
Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.
b11341 pinning
Added
- Pinned weights in
llama-server: when a MoE model is bigger than VRAM but fits in the RAM available at startup, the server loads its weights into pinned (page-locked) RAM instead of memory-mapping the file. The GPUs then read experts by direct DMA: 12k-token prompt processing GLM-5.3-Flash 179 -> 261 t/s (+46%, now 1.20x stock), Qwen3.8-Flash-Next 372 -> 435 t/s (+17%); decode is unchanged. Costs ~60 s more at startup (GLM: ~100 s vs 38 s) and keeps the model's RAM locked while the server runs, so it pays back after a few long prompts.llama-cliand the other tools keep mmap. --load-mode pinpins in any tool;--load-mode mmapkeeps mmap in the server.- If pinning leaves too little RAM (>2 GiB swapped while loading or <512 MiB left), the server reloads with mmap and logs why.
tools/moe-bench/: the frozen 12k/128k prompts,run.pyandperf.pybehind the README numbers, so anyone can reproduce them.
Changed
- Pinned memory is registered with
cudaHostRegisterinstead ofcudaMallocHost(about 2x faster to pin), and the loader drops the file's page cache behind it, so the file and its pinned copy never both sit in RAM. - The expert-cache upload log says whether it uploads from pinned or pageable memory.
Fixed
- Nothing.
Removed
- Nothing.
Regression check vs b11327 (same inputs, model in RAM)
| b11327 (mmap) | this release | |
|---|---|---|
| GLM tetris decode | 29.5 | 29.7 |
| GLM chat decode (avg of 2) | 20.4 | 20.5 |
| GLM 12k prompt / decode | 180 / 15.6 | 261 / 16.0 (pinned) |
| Qwen 12k prompt / decode | 372 / 38.4 | 435 / 39.0 (pinned) |
The pinned 12k numbers were measured on test builds with the same pinning code; this release was checked to pin (0 GiB swapped, 12.4 GiB RAM left with GLM). Remaining pinned + MTP numbers will be added to the README as they are measured.
Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.
b11327 qwen3next-mtp
Added
- Qwen3.8-Flash-Next MTP (PR #28243 by @danielhanchen), working together with the expert cache: 57.0 t/s with MTP vs 47.1 without (1.21x) on a 1500-token chat reply (2x RTX 3090, model already in RAM).
The MTP head is the small
llama-server -m Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf -md mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf --spec-type draft-mtpMTP/file from unsloth/Qwen3.8-Flash-Next-GGUF. - GLM-5.3-Flash MTP (PR #27917 by timkhronos), with GLM's MTP head loadable as a separate 4.3 GiB file: neuralll/GLM-5.3-Flash-MTP-GGUF, for any
glm5-nextGLM-5.3-Flash GGUF, no second copy of the model. - Automatic draft depth from measured speed. Each drafted token must be verified, and experts that miss the VRAM cache run on the CPU at full cost per drafted token, so the best depth depends on how much of the model fits in VRAM. The fork starts from a guess (model size vs free VRAM), then measures generation speed and moves the depth up or down.
LLAMA_SPEC_DEPTH=Npins it for benchmarks. - Draft skipped where it would be slower, with a warning: GLM-5.3-Flash at 2.3x VRAM on this box decodes 20.2 t/s without MTP vs 17.6 with it.
--spec-draft-n-max Nforces it on. With 3+ GPUs GLM should gain like Qwen.
Changed
- Default max draft depth: 2 when the model is bigger than free VRAM (the recurrent-state rollback buffers of hybrid models are sized by it and take cache VRAM), 5 when it fits; was 3.
--spec-draft-n-maxstill sets it. - The expert cache starts on the first real request instead of at context creation, so it sizes itself after a draft model or mmproj is loaded.
- Only the main context builds, steps and frees the expert cache; a draft context never touches it.
- Draft contexts use at most
-ub 512(drafts are a few tokens; 2048 only took cache VRAM).
Fixed
- Loading a draft model (
-md) with the expert cache failed with "unable to allocate CUDA buffer".
Removed
- Nothing.
Regression check vs b11297 (same inputs, model in RAM)
| b11297 | this release | |
|---|---|---|
| GLM tetris decode | 28.5 | 29.6 |
| GLM chat decode (avg of 2) | 20.4 | 20.4 |
| GLM 12k prompt / decode | 180 / 15.0 | 179 / 15.2 |
| Qwen chat decode, no MTP (avg of 2-3) | 47.2 | 47.1 |
| Qwen 12k prompt / decode (avg of 2) | 341 / 39.4 | 372 / 38.4 |
| Qwen chat decode, with MTP (automatic) | - (no MTP support) | 57.0 |
Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.
b11297 auto-defaults
Added
- Automatic defaults that differ from stock llama.cpp, only when a MoE model's weights are larger than the free VRAM of all GPUs (other models and every setting you pass behave as in stock).
llama-server -m model.ggufis enough; the chosen values are printed at startup (-lv 5prints all).
| setting | stock default | fork default | why |
|---|---|---|---|
| expert placement | autofit: whole layers on GPU, rest on CPU | all experts in RAM (--cpu-moe) |
the VRAM is used for the expert cache instead of fixed layers |
| expert cache | off | fills each GPU's free VRAM (--moe-expert-cache -1) |
holds the experts tokens actually use, so most expert work runs on the GPUs |
| CPU weight repacking | on | off (-nr) |
repacked experts can't be copied to the GPU cache |
-ub (tokens per prompt step) |
512 | 2048 if the largest GPU has 20+ GiB free, 1024 at 10+ GiB, else 512 | every expert upload serves more prompt tokens: ~2x faster long prompts; costs ~0.7 GiB cache VRAM (~1% decode) on GLM-5.3-Flash |
-c (context) |
model maximum, fitted to VRAM | 32768 | a bigger KV cache would take VRAM from the expert cache; -c 65536 if you need more |
-t (threads) |
all physical cores | cores minus one per GPU | a free core per GPU keeps kernel launches and cache uploads fast; measured faster |
- Prompt processing goes to the GPU in the fastest PCIe slot, on Linux and Windows: the fork measures each GPU's upload bandwidth at startup.
LLAMA_MOE_CACHE_DETERMINISTIC=1for reproducible benchmarks.- Lookahead expert prefetch from PR #28414 by @leshchukandrej (
--prefetch-experts-slots N, off by default).
Changed
- Decode is faster: the cache is published every step (+16% decode on long prompts), the VRAM safety margin dropped from 1 GB to 384 MiB per GPU, and small batches (up to 31 tokens: short prompts, speculative verification) go through the cache.
- Layers keep bus order (moving them to the fastest GPU cost ~7% decode on this box; see the README research notes).
- Latest upstream llama.cpp merged.
Fixed
- Qwen3.8-Flash-Next segfault on the first request with the cache.
- MiMo-V2.6 startup assert with the cache; MiMo no longer needs
-fitt.
Removed
- Nothing.
Decode: 1.5x to 2.5x stock on MoE models bigger than VRAM
2x RTX 3090 (one AM4 CPU PCIe 4.0 x16, one X570 chipset x4 slot) + Ryzen 7 3700X + 125 GB DDR4, stock llama.cpp -> this release:
| model | short prompt, 1500-token chat reply | 12k-token code prompt, decode after it |
|---|---|---|
| GLM-5.3-Flash 3.0-bit Q4_K attn (106 GB) | 13.8 -> 21.5 t/s (1.56x)* | 12.3 -> 15.7 t/s (1.28x)* |
| MiMo-V2.6-Flash-RL IQ3_XXS (132 GB) | 4.0 -> 10.1 t/s (2.54x)** | 4.2 -> 8.3 t/s (1.98x)** |
| Qwen3.8-Flash-Next UD-IQ4_XS (88 GB) | 27.7 -> 46.4 t/s (1.68x)* | 25.3 -> 38.7 t/s (1.53x)* |
* Model already in RAM (OS page cache), as on a server after its first request. The first run after switching to another large model is slower, once, while the file is read from disk. ** MiMo (132 GB) can't fully stay cached in 125 GB RAM.
Models that fit in VRAM (Qwen3.8-27B, OLMoE) are unaffected: same speed, identical output.
Prompt processing is still up to ~22% slower than stock (stock keeps whole layers in VRAM); using both GPUs' PCIe links for it is the next milestone.
Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.
b11214 expert-cache
Portable Linux CUDA build (GLM-5.3-Flash + VRAM-filling MoE expert cache fork). See README for benchmarks.
Important for prompt processing (prefill) speed: put the GPU in your fastest PCIe slot first. Prefill uploads the experts to the first GPU, and GPUs are numbered by PCI bus order, not slot speed, so the first one is often a card in a slow chipset x4 slot. This binary does not reorder automatically; set it yourself, e.g. CUDA_VISIBLE_DEVICES=1,0 llama-server ... (Linux) or set CUDA_VISIBLE_DEVICES=1,0 (Windows) when GPU 1 is the x16 card. Check with nvidia-smi --query-gpu=index,pcie.link.width.current --format=csv under load. On 2x RTX 3090 (x4 + x16) this took a 12k-token prompt from 45.9 to 84.1 t/s, and to 227.7 t/s adding -ub 2048 -b 2048.