What’s your preferred model? #28
Replies: 1 comment 1 reply
|
Hi, MOSS-TTS is the best open-weights model at the moment when memory or inference time is not a concern, IMO. It's expressive without sounding awkward, has the best voice likeness IMO, and has a nice, fluid delivery. Its accuracy is very good, though most all models released in the past half-year are all good with accuracy. I did play with LLM "auto-tagging" a little bit, but results so far have been... not that awesome, at least for the "bulk audio generation" use case. I added an experimental stand-alone script in The other model supporting a broad array of tags that I tried this with is Fish S2, but I found its tag behavior to be quite weak and inconsistent. |
Uh oh!
There was an error while loading. Please reload this page.
Just starting to play around with this and don't have time to try out all the models.
Do you have a recommendation for a tts model for the best overall quality? I have the hardware to run larger ones, so that's not a concern. Just looking for the best quality with, ideally one I can add emotional tags to.
Also, have you had any luck getting an LLM tag a book for emotional tags for a tts model to read?
All reactions