Skip to content

voices-v2 β€” fp32 piper + kokoro models (no fuzz)

Choose a tag to compare

@jphein jphein released this 07 May 03:23

Full-precision (fp32) versions of all voices in voices-v1.

The voices-v1 release used INT8-quantized ONNX models, which produced audible quantization fuzz on Samsung tablets (issue techempower-org/candela#6). These fp32 weights are pulled directly from k2-fsa/sherpa-onnx upstream tarballs and re-hosted as flat single-file downloads (one .onnx + one .tokens.txt per voice, plus kokoro-model.onnx + kokoro-voices.bin + kokoro-tokens.txt for the shared multi-speaker Kokoro model).

Naming: {lang}-{voice}-{quality}.onnx (e.g. en_US-lessac-high.onnx). voices-v1 had a latent collision where en_GB and en_US miro-high pointed at the same file β€” voices-v2 keeps them distinct via the language prefix.

Source: https://github.com/k2-fsa/sherpa-onnx/releases/tag/tts-models