What's new in 3.3.0 (2026-08-30)
These are the changes in inference v3.3.0.
New features
- feat(rerank): support jina-reranker-m0 by @Minamiyama in #5327
- feat(webui): show downloaded file sizes by @Minamiyama in #5330
- feat(model): support qwen3.8-27B by @bluefish-08 in #5337
- feat(webui): add IndexTTS emotion controls by @Minamiyama in #5335
- feat(audio): support multiple engines and MLX models by @qinxuye in #5334
- feat: [new model] minimax music3 by @Minamiyama in #5345
- feat(model): complete Qwen3.8 support by @qinxuye in #5339
- feat: support MiniMax-H3 Lightning LoRA by @qinxuye in #5338
- feat(audio): support FireRedTTS3 by @Minamiyama in #5352
- feat(llm): register DeepSeek V4 Flash 0731 by @m199369309 in #5371
- feat(image): add Ideogram4 support by @Minamiyama in #5367
- feat(image): sensenova u1.5 by @Minamiyama in #5369
- feat(image): support HiDream-O1 models by @Minamiyama in #5370
- feat(image): add GLM-Image support by @Minamiyama in #5394
- feat(router): add token-aware virtual model routing by @m199369309 in #5375
- feat(image): update SenseNova U1.5 model info by @Minamiyama in #5401
- feat(router): add tokenizer asset registry by @m199369309 in #5376
- feat: add media seed controls and dimension swap by @Minamiyama in #5400
- feat(router): bundle DeepSeek V4 tokenizer asset by @m199369309 in #5377
- feat(ui): add Token Router management by @m199369309 in #5378
- feat(router): add Agent-based Router orchestration by @m199369309 in #5379
- feat(llm): support Ornith-1.5-35B-A3B by @m199369309 in #5406
- feat(monitoring): add Token Router observability by @m199369309 in #5381
- feat(ui): add Router Agent and cluster management by @m199369309 in #5380
- feat(api): add unified Anthropic Messages protocol adapter by @m199369309 in #5389
- FEAT: [model] glm-5.2 support by @llyycchhee in #5404
- feat(venv): add secure per-launch find-links support by @m199369309 in #5387
- feat(router): route Anthropic Messages through virtual models by @m199369309 in #5391
- feat(router): expose virtual model running details by @m199369309 in #5397
- feat(venv): install Jina flash-attn wheel from find-links by @m199369309 in #5392
- feat(ui): add per-launch virtualenv find-links by @m199369309 in #5388
- feat(llm): support Ornith-1.5-397B by @m199369309 in #5405
- feat(audio): add ACE-Step 1.5 music generation support by @Minamiyama in #5413
- feat(image): add Krea 2 models by @Minamiyama in #5419
- FEAT: add world model generation support by @qinxuye in #5414
- FEAT: support replica config in launch CLI and Web UI by @m199369309 in #5422
- FEAT: [model] Kimi-K3 support by @llyycchhee in #5417
- FEAT: support configurable dynamic replica scale-up by @leslie2046 in #5426
- feat: add persistent system settings management by @Minamiyama in #5427
- FEAT: support NaviDC-OCR with Transformers and vLLM by @Minamiyama in #5431
- feat(embedding): add WeMM-Embedding support by @Minamiyama in #5439
- FEAT(model): Add Breeze-TTS-2 support by @Minamiyama in #5437
- feat(ui): show full paths on hover by @Minamiyama in #5444
Enhancements
- ENH: update models JSON [rerank, video] by @XprobeBot in #5328
- ENH(ui): make image previews adaptive by @Minamiyama in #5385
- Enh: Expose VoiceDesign ability to Qwen3-TTS-Voice-Design by @Minamiyama in #5355
- ENH: Add multi-engine video support with MLX by @qinxuye in #5407
- ENH(UI): Add configurable speech and music output formats by @Minamiyama in #5436
- ENH: support Base64 video input in vLLM chat by @amumu96 in #5434
Bug fixes
- fix(audio): correct IndexTTS-2.5 runtime dependencies by @Minamiyama in #5331
- fix(audio): return playable speech responses by @Minamiyama in #5332
- BUG: pin DeepDoc transformers to >=4.51,<5 for OCR launches by @1084669403 in #5340
- fix(ui): auto-size launch model dropdowns by @Minamiyama in #5343
- fix(core): guard missing/invalid usage in non-stream chat metrics by @Hyhyhyyy in #5342
- fix(model): calculate download progress per file by @Minamiyama in #5346
- fix(ui): include scrollbar in dropdown width by @Minamiyama in #5354
- fix: validate LLM metadata before packaging by @li2631026381-alt in #5353
- fix(core): refine autostart transition handling and reset retry attempts by @baka-world in #5361
- fix(worker): wait reliably for metrics exporter startup by @baka-world in #5363
- fix(core): make worker supervisor reconnection generation-safe by @baka-world in #5364
- fix(model): harden batch processor against cancellation and unexpected exits by @baka-world in #5362
- fix(core): normalize distributed worker count by @baka-world in #5366
- fix(llm): preserve DeepSeek V4 tool argument types by @m199369309 in #5372
- fix(vllm): handle engine death during async iteration by @m199369309 in #5374
- fix(vllm): use block size 256 for DeepSeek V4 by @m199369309 in #5373
- fix(api): map Qwen3.8 top-level reasoning effort by @baka-world in #5383
- fix(image): pin Qwen Layered dependencies by @Minamiyama in #5390
- fix(cache): fully remove downloaded models by @Minamiyama in #5395
- fix(supervisor): reschedule autostart when the last replica dies by @AmirF194 in #5409
- fix(venv): honor configured sources in post-install hooks by @m199369309 in #5386
- fix(ui): clarify Token Router identifiers and asset layout by @m199369309 in #5393
- fix(router): separate runtime and backend credentials by @m199369309 in #5396
- fix(router): stabilize managed runtime instance identity by @m199369309 in #5398
- fix(ui): harden Token Router running model and detail views by @m199369309 in #5399
- fix(llm): support Hugging Face cache roots in auto-fill by @amumu96 in #5403
- FIX: stabilize MLX runtime dependencies and locks by @qinxuye in #5423
- fix: declare Token Router installation extra by @m199369309 in #5424
- fix(supervisor): reschedule autostart on whole-worker death by @AmirF194 in #5429
- fix: skip cache status update when cache tracker ref is None by @BetterAndBetterII in #5443
- FIX: preserve available model download hubs by @Minamiyama in #5438
- fix(image): classify OvisOCR2 as OCR by @Minamiyama in #5435
- FIX: stabilize HY-WorldPlay runtime and progress reporting by @qinxuye in #5432
- FIX: restore GPU and aarch64 Docker builds by @qinxuye in #5451
- FIX: repair pypiserver platform dependencies by @qinxuye in #5453
Others
- docs: add v3.2.0 to release notes index and locale catalogs by @martinma51 in #5348
- ci: skip GPU tests for video and flexible model changes by @OliverBryant in #5368
- refactor(router): remove legacy Agent compatibility by @m199369309 in #5421
- chore: refresh built-in model catalog and skip draft CI by @qinxuye in #5411
New Contributors
- @1084669403 made their first contribution in #5340
- @Hyhyhyyy made their first contribution in #5342
- @li2631026381-alt made their first contribution in #5353
- @baka-world made their first contribution in #5361
- @AmirF194 made their first contribution in #5409
- @BetterAndBetterII made their first contribution in #5443
Full Changelog: v3.2.0...v3.3.0