QAI Hub Models v0.60.0 — Release Notes (2026-08-10)
✨ New Models
- Whisper Large v3 Turbo (quantized) (
whisper_large_v3_turbo_quantized) — new quantized speech-to-text model - OWL-ViT (
owl_vit) — new open-vocabulary object detection model
🧠 LLMs
- Gemma-4 E4B and Qwen3-0.6B are now supported via the GenieX QAIRT plugin
Note: Gemma-4 E4B is still under Developer Preview. We're working to further improve on-device quantization and
accuracy. Known issues include repetitive words, formatting glitches, and code-related questions. For more accurate responses,
we recommend the GenieX llama.cpp plugin-based model (quantization techniqueq4_0). - Qwen3-1.7B — ~15% faster on-device with no accuracy loss; improved response quality on-device
- Qwen3-0.6B — improved accuracy on the same-size checkpoint; now ships with
[512, 1024, 4096]context-length variants - Fixed Qwen3-8B floating-point model that produced incorrect outputs via
evaluate/demo - Fixed regression in LLM export via the
qai-hub-modelsCLI
🔧 Model Improvements & Fixes
- PT2 export is now the default — most models automatically use
torch.export.export - Improved quantized accuracy for:
- PoseNet-MobileNet
- DETR ResNet-50 DC5
- EfficientViT
- UNet Segmentation
- LeViT
- MobileNet v3 Large
- Added evaluation support for OpenAI CLIP and DeepBox
💻 CLI
- New command
qai-hub-models install <model>— simpler way to install a model's dependencies - Model name and chipset arguments now support fuzzy matching
📊 Performance
- All performance numbers updated using the new PT2 recipe
- Added performance numbers for new device Ventuno Q