New Model Support
- Gemma 3n (#1650 by @anatyrova)
- Qwen3-Omni-MOE (#1700 by @sgonorov)
- SmolLM3 (#1715 by @AshutoshSinghIntel)
- Qwen3-VL-Embedding (#1723 by @mlukasze)
- Gemma 4 Unified (#1770 by @rkazants)
- FLUX.2 (#1809 by @rkazants)
Improvements & fixes
- Transformers v5.1-5.5 support (#1684 by @echarlaix)
- Fix Kokoro voice conversion from local model directory (#1803 by @mmikolajcz)
- Extended OpenVINO transformation verification tests for MoE architectures (#1776 by @anatyrova)
- Unpatch 16-bit models previously patched with
patch_16bit_modelafter conversion (#1811 by @anatyrova) - Set dynamic sequence length for FLUX.2 (#1846 by @rkazants)
- Deprecated the broken
contextualcalibration dataset, replaced withtextvqa(#1849 by @echarlaix) - Fixed passing of
token_type_idsfor Gemma3 (#1881 by @popovaan) - Fixed export for encoder-decoder models with unnamed tensors (#1857 by @echarlaix)
- Fixed long audio inference for Qwen3-ASR (#1782 by @popovaan)
- Fixed ssm states handling for Mamba and Falcon-Mamba (#1779 by @echarlaix)
- Pinned
huggingface_hub<1.22to fix offline loading (#1855 by @echarlaix) - Use base transformers version for version comparisons (#1798 by @AlexanderDokuchaev)
Other Changes
- CI: local model caching and rate-limit fixes for fork PRs (#1817, #1826, #1835, #1837, #1840, #1842 by @echarlaix)
- Faster doc builds (#1851 by @mishig25)
New Contributors
- @anatyrova made their first contribution in #1650
- @AshutoshSinghIntel made their first contribution in #1715
- @mlukasze made their first contribution in #1723
Full Changelog
Compatible transformers version
Compatible with transformers>=4.51,<5.6
NPU-specific requirements:
- Whisper - need transformers v5.0
- LFM2 - need transformers v5.0
Maximum supported transformers version by model architectures can be found in the table below (models not listed support up to transformers v5.6)
| Model architecture | Max transformers version |
|---|---|
| AquilaM | 4.57.6 |
| Arctic | 4.53.3 |
| Baichuan | 4.57.6 |
| Bitnet | 4.57.6 |
| ChatGLM2 | 4.55.4 |
| DBRX | 4.57.6 |
| Data2VecText | 4.57.6 |
| Deci | 4.57.6 |
| Deepseek | 4.53.3 |
| Exaone | 4.57.6 |
| Exaone4 | 4.57.6 |
| Flaubert | 4.57.6 |
| Gemma | 5.0 |
| Gemma2 | 5.0 |
| Gemma3Text | 5.0 |
| Gemma3nText | 5.0 |
| Gemma4Text | 5.0 |
| Gemma4Unified | 5.10.4 |
| Gemma4UnifiedText | 5.10.4 |
| GLM | 5.0 |
| GotOCR2 | 4.57.6 |
| GraniteMoeHybrid | 5.3.0 |
| Idefics3 | 4.57.6 |
| InternLM | 4.57.6 |
| InternLM2 | 4.57.6 |
| InternVLChat | 4.57.6 |
| Jais | 4.57.6 |
| LFM2 | 5.4.0 |
| LFM2Moe | 5.4.0 |
| Llama4 | 4.57.6 |
| Llama4Text | 4.57.6 |
| LlavaNextVideo | 4.57.6 |
| LlavaQwen2 | 4.53.3 |
| Mamba | 5.3.0 |
| Marian | 4.57.6 |
| MiniCPM | 4.53.3 |
| MiniCPM3 | 4.53.3 |
| MiniCPMO | 4.51.3 |
| MiniCPMV | 4.57.6 |
| MT5 | 4.57.6 |
| Nystromformer | 4.50.3 |
| Orion | 4.57.6 |
| Phi3Vision | 4.53.3 |
| Phi4MM | 4.53.3 |
| Qwen | 4.55.4 |
| Qwen2VL | 5.0 |
| Qwen2_5_VL | 5.0 |
| Qwen3ASR | 4.57.6 |
| Qwen3Next | 4.57.6 |
| Qwen3VL | 5.0 |
| Qwen3_5 | 5.2.0 |
| Qwen3_5Moe | 5.2.0 |
| Qwen3_5MoeText | 5.2.0 |
| Qwen3_5Text | 5.2.0 |
| SmolVLM | 4.57.6 |
| VideoChatFlashQwen | 4.57.6 |
| XLM | 4.57.6 |
| XverseM | 4.57.6 |
| Zamba2 | 4.57.6 |
Recommended versions
- OpenVINO: v2026.3
- OpenVINO GenAI: v2026.3
- NNCF: v3.3