Repository navigation
v3.1.0
·
3073 commits
to develop
since this release
- General:
- Migrated
NNCFGraphfromnx.DiGraphtonx.MultiDiGraphto support models with parallel/multi-edges, enabling correct quantization of models with complex graph structures such as YOLO26 and models likea = conv(x); return a * a(#3843).
- Migrated
- Features:
- (OpenVINO) Added NVFP4 (
f4e2m1) compression data type in Weight Compression. NVFP4 uses a constant group size of 16 with scales compressed tof8e4m3using a second-degree scale (#3967). - (OpenVINO) Added
backup_modeparameter for FP compression formats (MXFP4, MXFP8, FP4, FP8), allowing first/last layers to be compressed with a backup FP format instead of INT8 (#3886). - (OpenVINO) RoPe ignored pattern is updated to handle operations without a preceding transpose like in the Phi-3.5-MoE-instruct model (#3989).
- (PyTorch) Added
TopKMetatypesupport for TorchFX backend, enabling correct graph building for models with TopK operations such as YOLO26 (#3944). - (PyTorch) Migrated to use
torchaoinstead of deprecatedtorch.ao(#3854).
- (OpenVINO) Added NVFP4 (
- Fixes:
- Improvements:
- (PyTorch) Added lazy import for
nncf.torchmodule to reduce startup import time (#3862).
- (PyTorch) Added lazy import for
- Tutorials:
- Post-Training Optimization of Gemma 4 Model
- Post-Training Optimization of Code-specialized LLMs
- Post-Training Optimization of Vision-Language Models (VLMs)
- Post-Training Optimization of MiniCPM-o 4.5 Multimodal Model
- Post-Training Optimization of PaddleOCR-VL/PaddleOCR-VL-1.5 Models
- Post-Training Optimization of RAG pipeline
- Requirements: