Skip to content

v0.17.7 — Paradee-8M

Latest

Choose a tag to compare

@Alex-Wengg Alex-Wengg released this 08 Oct 02:48
· 3 commits to main since this release

Paradee-8M (beta) — ParadeeManager (FluidInference/paradee-8m-coreml). Sahil Mahendrakar's Paradee distills Kokoro-82M into an 8M-parameter, single-voice (af_heart) American English TTS model. The Core ML build is 12 MB (int8, default) or 34 MB (fp32) and reuses the Kokoro English frontend.

MiniMax English, 100 phrases, M5 Pro, CPU Paradee int8 Paradee fp32
Download 12 MB 34 MB
WER / CER (Parakeet round trip) 1.20 % / 0.14 % 1.20 % / 0.14 %
RTFx (incl. G2P) ~100× ~100×
Median synth time per phrase 77 ms 77 ms

On the same inputs, the Core ML models run ~150× real time against ~20× for the upstream ONNX on one thread (~47× on all 15 threads), with identical Whisper transcripts.

let manager = ParadeeManager()
try await manager.initialize()
let samples = try await manager.synthesize(text: "Hello from FluidAudio.")  // 24 kHz Float32

CLI: tts --backend paradee [--variant int8|fp32] [--speed] [--seed] [--phonemes]; tts-benchmark --backend paradee. Runs on .cpuOnly (default) or .cpuAndNeuralEngine; GPU compute units are rejected (the LSTMs abort in MPSGraph). See Paradee.md.

Also in this release: Kokoro ANE English now applies misaki's ɾ → T flap step, so words like "metal" no longer come out as "mattle" (#999), and updated Kokoro v3 benchmark, compute-routing and ANE-profiling docs (#1000).

What's Changed

  • feat(tts): Paradee-8M CoreML backend [beta] by @Alex-Wengg in #998
  • fix(tts/kokoro-ane): apply misaki's ɾ → T flap step to English phonemes by @Alex-Wengg in #999
  • docs(kokoro): update v3 benchmarks, compute routing and ANE profiling in #1000

Full Changelog: v0.17.6...v0.17.7