Skip to content

b11367 mtp-fixes

Choose a tag to compare

@neurall neurall released this 26 Sep 17:41
· 34 commits to release since this release

Fixed

  • MTP draft-depth tuner picked the wrong depth by chance. Early in a reply (the model's reasoning) depths measure almost tied, so noise decided; on Qwen3.8-Flash-Next it sometimes settled on depth 1 (~51 t/s) instead of 2 (~60 t/s). The tuner now keeps its current depth (at first the model-size guess) unless another depth is at least 5% faster. Qwen3.8-Flash-Next + MTP chat: 60.4 t/s (2.18x stock), with pinned weights (prompt processing keeps the pinned +45%).

Changed

  • Built-in MTP layers are no longer skipped on models much bigger than VRAM. The ">2x VRAM: don't load a draft" rule now applies only to a separate -md draft (GLM-5.3-Flash's costs the expert cache more than it gains). Built-in MTP, e.g. MiMo-V2.6's 3 dense MTP layers, starts at depth 0 and the tuner decides; on MiMo-V2.6-Flash-RL here it keeps depth 0 (drafting was not 5% faster), so no measurable gain on this box.
  • tools/moe-bench/perf.py: a discarded run before every measured run (model switches and a previous pinned load caused 5-12% swings on first runs); leftover servers are killed by process name.

Added

  • tools/moe-bench/perf.db: every benchmark run behind the README numbers.
  • tools/moe-bench/glm_splice_mtp.py / gguf_remote.py: how the GLM-5.3-Flash MTP head on Hugging Face was made (the MTP tensors are pulled from unsloth's GGUF with HTTP range requests, ~4.3 GiB instead of the whole model).

Removed

  • Nothing.

Regression check vs b11341

Only the MTP depth tuner and defaults changed (no kernels or cache code):

b11341 this release
Qwen chat + MTP head (-md, default settings) ~51 (tuner could pick depth 1) 60.4 (depth 2; 4 more runs: 55.8-62.3, all depth 2)
MiMo chat, --spec-type draft-mtp MTP skipped loaded, tuner keeps depth 0 (no gain)

GLM, MiMo and 12k MTP numbers are being re-measured with the fixed tuner and will be updated in the README.

Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.