What's Changed
- vendor Gemma/Phi/GLM/Cohere text models and reuse shared MLPs by @Lazarus-931 in #1616
- fix(apc): exact hybrid + TurboQuant warm match live kv layout by @weklund in #1596
- vendor 24 more text models as native packages by @Lazarus-931 in #1625
- Speed up llguidance structured decoding on Qwen 3.5 by @Blaizzy in #1628
- Improve server generation logging by @Blaizzy in #1634
- Fix LFM2-VL tokenizer loading without PyTorch by @Blaizzy in #1640
- Fix streaming reasoning classification by @Blaizzy in #1641
- Fix streaming reasoning protocol compatibility by @Blaizzy in #1644
- Bump version to 0.6.6 by @Blaizzy in #1645
- Fix Qwen3-VL PIL video inputs by @Blaizzy in #1642
Full Changelog: v0.6.5...v0.6.6