voxcpm2-4bit @ main
·
54 commits
to main
since this release
Automated model bundle release for voxcpm2-4bit.
Description: Multilingual MLX text-to-speech and voice cloning model for local assistant voices.
Use cases
- Spoken assistant responses
- Voice cloning with short reference audio and transcript
- Local TTS playback in Flutter chat
- Voice UX prototyping
Limitations
- Voice controls vary by model family
- Generated speech quality depends on prompt and runtime support
- Use a clean 3-10 second reference clip and transcript for stable voice cloning
- Requires a GitHub mlx-audio build because PyPI mlx-audio 0.4.3 does not include VoxCPM2 yet
Release metadata
- Runtime:
mlx_audio - Tasks:
text_to_speech, audio_output - Source:
mlx-community/VoxCPM2-4bit@main - Archive size:
2302095360bytes - Languages:
multilingual - Output:
{"media_type": "audio/wav", "extension": "wav", "supports_inline_playback": true} - Default parameters:
{"audio_format": "wav", "join_audio": true, "lang_code": "en", "cfg_scale": 2.0, "ddpm_steps": 7, "max_tokens": 2000} - Voices:
Default VoxCPM2
This release includes model_metadata.json, release_metadata.json, manifest.source.yaml, and chunked model archive assets.