1.4.5
- Support Glm5NextForConditionalGeneration (GLM5.3-Flash)
- Support Qwen4ExpForConditionalGeneration (Qwen3.8-Flash-Next)
- Improved MoE MTP performance
- Slab allocation for MoE layers to avoid wasted VRAM for badly aligned expert tensor shapes
- Work around Triton race condition (
autotuner: 'NoneType' object is not a mappingerror) - Fix Unigram tokenizers
- Other bugfixes and QoL improvements
Full Changelog: v1.4.4...v1.4.5