1.4.3
- Support GlmMoeDsaForCauslLM (GLM 5.2, still probably somewhat WIP)
- Partial CPU layer expert offloading option with dynamic placement (supersedes expert cache)
- Improved dynamic draft sizing with auto-calibrated confidence thresholds
- New (experimental) quant-optimizer pipeline
- Support for mid-stream text injection in generator (enables reasoning token budget)
- More precise autosplit allocation
- Reduced (and now stable) VRAM footprint for DSA prefill
- CPU cache offloading supported in TP mode
- Bugfixes and QoL improvements
- Remove incomplete Nanochat implementation
Full Changelog: v1.4.2...v1.4.3