b11367 mtp-fixes
·
34 commits
to release
since this release
Fixed
- MTP draft-depth tuner picked the wrong depth by chance. Early in a reply (the model's reasoning) depths measure almost tied, so noise decided; on Qwen3.8-Flash-Next it sometimes settled on depth 1 (~51 t/s) instead of 2 (~60 t/s). The tuner now keeps its current depth (at first the model-size guess) unless another depth is at least 5% faster. Qwen3.8-Flash-Next + MTP chat: 60.4 t/s (2.18x stock), with pinned weights (prompt processing keeps the pinned +45%).
Changed
- Built-in MTP layers are no longer skipped on models much bigger than VRAM. The ">2x VRAM: don't load a draft" rule now applies only to a separate
-mddraft (GLM-5.3-Flash's costs the expert cache more than it gains). Built-in MTP, e.g. MiMo-V2.6's 3 dense MTP layers, starts at depth 0 and the tuner decides; on MiMo-V2.6-Flash-RL here it keeps depth 0 (drafting was not 5% faster), so no measurable gain on this box. tools/moe-bench/perf.py: a discarded run before every measured run (model switches and a previous pinned load caused 5-12% swings on first runs); leftover servers are killed by process name.
Added
tools/moe-bench/perf.db: every benchmark run behind the README numbers.tools/moe-bench/glm_splice_mtp.py/gguf_remote.py: how the GLM-5.3-Flash MTP head on Hugging Face was made (the MTP tensors are pulled from unsloth's GGUF with HTTP range requests, ~4.3 GiB instead of the whole model).
Removed
- Nothing.
Regression check vs b11341
Only the MTP depth tuner and defaults changed (no kernels or cache code):
| b11341 | this release | |
|---|---|---|
Qwen chat + MTP head (-md, default settings) |
~51 (tuner could pick depth 1) | 60.4 (depth 2; 4 more runs: 55.8-62.3, all depth 2) |
MiMo chat, --spec-type draft-mtp |
MTP skipped | loaded, tuner keeps depth 0 (no gain) |
GLM, MiMo and 12k MTP numbers are being re-measured with the fixed tuner and will be updated in the README.
Linux and Windows CUDA builds; on Windows also extract the cudart zip if you don't have CUDA 13.4 installed.