v0.0.22
DeepFilterNet3 quality
DeepFilterNet3 speech enhancement now matches the reference libdf DSP conventions while retaining the compact Core ML model: 8-bit palettized weights with FP16 compute. On the documented VoiceBank-DEMAND 20-clip benchmark, the Swift path reaches PESQ-WB 3.097, STOI 0.964, and SI-SDR 19.12 dB.
What's Changed
- Create FUNDING.yml by @ivan-digital in #316
- Improve realtime websocket error handling by @ivan-digital in #315
- Fix funding configuration syntax in FUNDING.yml by @ivan-digital in #317
- Keep realtime websockets alive during cold starts by @ivan-digital in #318
- Keep realtime inference off websocket event loops by @ivan-digital in #320
- Disable Hummingbird auto-ping on realtime websocket by @ivan-digital in #321
- Lift realtime websocket keepalive to session scope by @ivan-digital in #322
- feat(qwen3-tts): ship 1.7B bf16, drop degrading 1.7B int4 by @ivan-digital in #324
- sortformer: high-throughput offline variant + proper streaming state (#319) by @ivan-digital in #326
- FunctionGemma — on-device tool-calling LLM (Gemma 3 270M, CoreML) by @ivan-digital in #327
- fix(sortformer): default to ANE-only compute units (#319) by @ivan-digital in #328
- feat(cosyvoice): honor style instructions on cloned voices via instruct2 by @ivan-digital in #323
- docs(tts): refresh roundtrip benchmark for the 1.7B int4 drop by @ivan-digital in #325
- feat(sortformer): preload(), computeUnits, balanced variant by @ivan-digital in #329
- fix(qwen3): correct 1.7B speaker-encoder dims and conv layout for voice cloning by @ivan-digital in #330
- fix(qwen3): pin voice-clone tests to one model to stop x-vector collapse by @ivan-digital in #332
- chore(tts): decommission int4 for TTS models (defaults → bf16/int8) by @ivan-digital in #333
- docs(tts): drop int4 from the TTS model catalog after the decommission by @ivan-digital in #334
- FunctionGemma: render full developer/declaration block via FunctionGemmaPrompt by @ivan-digital in #335
- chore(tts): point CustomVoice at the new bf16 bundle (int4 fully decommissioned) by @ivan-digital in #336
- feat(downloader): support HF_ENDPOINT mirror for users in China by @hanrw in #331
- fix(function-gemma): route loadFromHub through HF_ENDPOINT mirror by @ivan-digital in #337
- Qwen3-ASR: add 5-bit MLX quantization variant by @xzjh in #262
- skill(benchmark): /benchmark asr now reports WER + peakRSS + throughput by @ivan-digital in #338
- feat(asr): wire 5-bit MLX variants into AsrBenchmark + docs by @ivan-digital in #340
- Add SupertonicTTS-3 (CoreML) — non-autoregressive flow-matching TTS by @ivan-digital in #341
- fix(qwen3-icl): reimplement the Mimi codec encoder so voice clone is target-only by @ivan-digital in #339
- docs(readme): swap Product Hunt badge for Trendshift badge by @ivan-digital in #344
- feat(forced-aligner): CoreML runtime + speech align --engine coreml by @ivan-digital in #345
- ChatterboxTTS: multilingual voice cloning in Swift/MLX by @ivan-digital in #343
- Fix Qwen3.5 streaming decode corruption + empty replies by @ivan-digital in #346
- ChatterboxTTS: fix from-pretrained dropping model.safetensors by @ivan-digital in #347
- Add Qwen3 dense MLX runtime (Qwen3-4B chat) + quant/sampler fixes by @ivan-digital in #348
- OmniVoiceTTS: MLX-Swift port of OmniVoice (600+ language NAR diffusion TTS) by @ivan-digital in #349
- Fix OmniVoice instruct conditioning by @ivan-digital in #351
- Fix nightly E2E shard stability by @ivan-digital in #352
- Group README models by capability by @ivan-digital in #354
- Add daily benchmark dashboard by @ivan-digital in #353
- Add Gemma 4 chat backend by @ivan-digital in #355
- Add local benchmark runner by @ivan-digital in #357
- Use Silero VAD v6.2.1 CoreML by default by @ivan-digital in #359
- Add Hindi emotion TTS candidates by @ivan-digital in #358
- Add Indic-Mio raw reference cloning by @ivan-digital in #360
- Add Fish Audio S2 Pro runtime by @ivan-digital in #361
- Use Silero VAD v6.2.1 MLX by default by @ivan-digital in #362
- Split Kokoro fixture from Qwen ASR tests by @ivan-digital in #363
- Update model documentation coverage by @ivan-digital in #364
- Add resumable Fish Audio downloads by @ivan-digital in #365
- Speed up ranged Hugging Face downloads by @ivan-digital in #366
- Add WhisperASR CoreML port by @ivan-digital in #367
- Honor PersonaPlex demo offline cache by @ivan-digital in #370
- Speed up diarization: batch segmentation windows, memoize clustering distances by @JimLiu in #371
- Add communication guidelines and PR review skill by @ivan-digital in #373
- Fix streaming audio startup de-click by @ivan-digital in #374
- Support Chatterbox multilingual frontends by @ivan-digital in #375
- Add Audio2Face-3D avatar motion runtime by @ivan-digital in #376
- Skip CosyVoice LLM micro-test in nightly by @ivan-digital in #377
- CosyVoice: seeded long-form synthesis + upstream parity fixes by @ivan-digital in #378
- Fix Chatterbox CoreML int8 SDK compatibility by @ivan-digital in #380
- Add Chatterbox Flash voice cloning bridge by @ivan-digital in #381
- Update AGENTS.md to match current codebase by @rymalia in #379
- Clarify agent testing commands by @ivan-digital in #382
- Add native IndexTTS2 voice cloning runtime by @ivan-digital in #384
- Add voice cloning candidate runtimes by @ivan-digital in #383
- Add native F5-TTS voice cloning runtime by @ivan-digital in #385
- Add Mandarin pinyin frontend to F5-TTS by @ivan-digital in #386
- State current F5-TTS language support without legacy-bundle caveats by @ivan-digital in #387
- Add native Higgs TTS 3 voice cloning runtime by @ivan-digital in #389
- Pipeline the Higgs decode loop and tune voice-cloning TTS performance by @ivan-digital in #390
- Cache and batch the IndexTTS2 semantic GPT decode by @ivan-digital in #391
- Rewrite BigVGAN FIR resampling as dense convolution by @ivan-digital in #392
- Optimize Qwen3-ASR batch decode CPU sync by @hhh2210 in #234
- Cleanup follow-ups for Qwen3-ASR batched decode by @ivan-digital in #236
- Expose IndexTTS2 S2Mel flow steps by @ivan-digital in #393
- Add isolated E2E regression runner by @ivan-digital in #394
- Fix Magpie Chinese frontend to match NeMo tokenization by @ivan-digital in #395
- Fix CustomVoice: retire phantom 8bit bundle, stabilize float16 decode by @ivan-digital in #396
- Cut IndexTTS2 GPT step overhead, default S2Mel to 15 steps by @ivan-digital in #397
- Harden downloader against large-shard stalls and partial bundles by @ivan-digital in #398
- Finalize empty Qwen pipeline generations by @ivan-digital in #400
- Add Parakeet language steering by @ivan-digital in #399
- Eliminate O(history) KV copies in IndexTTS2 GPT decode by @ivan-digital in #401
- Run IndexTTS2 GPT attention in float32 to dodge fp16 SDPA slowdown by @ivan-digital in #402
- Add SystemAudioTap: system-output capture via Core Audio process taps by @ivan-digital in #403
- Skip system audio tap E2E on hosted nightly runner by @ivan-digital in #404
- Fix centroid mapping after speaker ID compaction by @ivan-digital in #405
- Add Community-1 CoreML diarization runtime by @ivan-digital in #408
- Match DeepFilterNet3 enhancement with reference libdf DSP by @ivan-digital in #407
New Contributors
- @hanrw made their first contribution in #331
- @xzjh made their first contribution in #262
- @JimLiu made their first contribution in #371
- @rymalia made their first contribution in #379
Full Changelog: v0.0.21...v0.0.22