- Highlights
- Features
- Improvements
- Validated Hardware
- Validated Configurations
Highlights
- NVFP4 with E5M3 scaling experimental support for PyTorch
- Selective and mixed quantization support for Keras/JAX
Features
- Support NVFP4 with E5M3 scaling for PyTorch (Experimental)
- Support model-free quantization for PyTorch
- Add dynamic fused MoE bias support for GPT-OSS model on Gaudi
- Add vLLM QDQ plugin for accuracy benchmark based on simulation
- Support for selective and mixed quantization in Keras/JAX, allowing different model components to remain unchanged or use different quantization methods and configurations, including static or dynamic quantization, different data types, and other quantization options
- Support for quantization per channel for weights for Keras/JAX
Improvements
- Security issue fixes
- SDXL model MXFP8 PTQ example
- Kimi K2.6 MXFP4 PTQ example
- Minimax-M2.7 MXFP4 PTQ example
Validated Hardware
- Intel Gaudi Al Accelerators (Gaudi 2 and 3)
- Intel Xeon Scalable processor (4th, 5th and 6th Gen)
- Intel® Arc™ B-Series Graphics GPU (B60)
Validated Configurations
- Ubuntu 24.04 & Win 11
- Python 3.11, 3.12, 3.13
- PyTorch 2.12, 2.13, 2.14
- JAX 0.10
Notes
- It is recommended to use version v3.10 or later to mitigate code CVEs.