Skip to content

v0.60.0

Latest

Choose a tag to compare

@qaihm-bot qaihm-bot released this 13 Aug 18:39
· 52 commits to main since this release
5347a3b

QAI Hub Models v0.60.0 — Release Notes (2026-08-10)

✨ New Models

  • Whisper Large v3 Turbo (quantized) (whisper_large_v3_turbo_quantized) — new quantized speech-to-text model
  • OWL-ViT (owl_vit) — new open-vocabulary object detection model

🧠 LLMs

  • Gemma-4 E4B and Qwen3-0.6B are now supported via the GenieX QAIRT plugin

    Note: Gemma-4 E4B is still under Developer Preview. We're working to further improve on-device quantization and
    accuracy. Known issues include repetitive words, formatting glitches, and code-related questions. For more accurate responses,
    we recommend the GenieX llama.cpp plugin-based model (quantization technique q4_0).

  • Qwen3-1.7B — ~15% faster on-device with no accuracy loss; improved response quality on-device
  • Qwen3-0.6B — improved accuracy on the same-size checkpoint; now ships with [512, 1024, 4096] context-length variants
  • Fixed Qwen3-8B floating-point model that produced incorrect outputs via evaluate / demo
  • Fixed regression in LLM export via the qai-hub-models CLI

🔧 Model Improvements & Fixes

  • PT2 export is now the default — most models automatically use torch.export.export
  • Improved quantized accuracy for:
    • PoseNet-MobileNet
    • DETR ResNet-50 DC5
    • EfficientViT
    • UNet Segmentation
    • LeViT
    • MobileNet v3 Large
  • Added evaluation support for OpenAI CLIP and DeepBox

💻 CLI

  • New command qai-hub-models install <model> — simpler way to install a model's dependencies
  • Model name and chipset arguments now support fuzzy matching

📊 Performance

  • All performance numbers updated using the new PT2 recipe
  • Added performance numbers for new device Ventuno Q