Hi team,
Thank you for the excellent work on OmniVoice — the support for 600+ languages and zero-shot voice cloning is truly impressive .
However, the current model is relatively heavy for on-device usage (e.g., ~4GB VRAM on a laptop), which limits real-time and edge deployment scenarios.
I wanted to ask if there are any plans to release a lightweight or distilled version of OmniVoice (e.g., ~100M–300M parameters), similar in spirit to ZipVoice, that:
Supports a smaller subset of languages (e.g., 5–10)
Retains voice cloning capability
Is optimized for edge / real-time inference
Such a model would be highly valuable for deployment with frameworks like Sherpa-ONNX / edge pipelines.
Looking forward to your thoughts.
Thanks!
Hi team,
Thank you for the excellent work on OmniVoice — the support for 600+ languages and zero-shot voice cloning is truly impressive .
However, the current model is relatively heavy for on-device usage (e.g., ~4GB VRAM on a laptop), which limits real-time and edge deployment scenarios.
I wanted to ask if there are any plans to release a lightweight or distilled version of OmniVoice (e.g., ~100M–300M parameters), similar in spirit to ZipVoice, that:
Supports a smaller subset of languages (e.g., 5–10)
Retains voice cloning capability
Is optimized for edge / real-time inference
Such a model would be highly valuable for deployment with frameworks like Sherpa-ONNX / edge pipelines.
Looking forward to your thoughts.
Thanks!