What's Changed
- Vendor standard models by @Lazarus-931 in #1702
- Fix Laguna XS NVFP4 sanitizer by @Blaizzy in #1667
- Add Mage-Flow image model by @lucasnewman in #1708
- Authenticate inference API routes by @Blaizzy in #1714
- Awq quant by @Lazarus-931 in #1666
- Support mixed-precision compressed-tensors checkpoints (fp8 + NVFP4) by @Lazarus-931 in #1711
- Fix Responses instructions for Qwen chat templates by @Blaizzy in #1716
- Return model loading failures as bad requests by @Blaizzy in #1717
- Allow image editing with FLUX.2-klein-4b by @lucasnewman in #1729
- gpt_oss: honor per-layer mixed quantization (mxfp4 experts + 8-bit affine) by @Lazarus-931 in #1728
- Support quantized FLUX.2 models by @lucasnewman in #1730
- Add PLaMo 2.1 VL (plamo2vl) model support by @takuyaomi in #1739
- Add llm-jp-4-vl (llmjpvl) model support by @takuyaomi in #1738
- Fix chunked MRoPE positions across Qwen VL models by @Blaizzy in #1741
- Fix multi-image SFT grid collation by @Blaizzy in #1744
New Contributors
- @takuyaomi made their first contribution in #1739
Full Changelog: v0.6.7...v0.6.8