- Fix correctness bug in split MoE CPU-offload mode
- Better MTP support in expert-parallel mode
- Validate and default to vision tower quantization
- Support vision model offloading (streams from system memory, performance penalty is small)
Full Changelog: v1.4.3...v1.4.4