Description
I am using a reference audio, cloned voice and the output for what should be approximately one minute of output is ALWAYS 41-48 seconds. I have tried a multitude of modifications based on the available parameters, including temperature, subtalker temperature, do_sample and subtalker_dosample. None of these parameters appear to have an effect on speaking rate.
Something that I did as a test was to change the input reference to 20% slower (using Audacity or other program), making the audio reference significantly slower than the original. Yet, the one minute of text outcome was still time compressed to 41-48 seconds and as you get closer to the end, pausing and stopping for punctuation is almost always greatly compressed.
The speaker pacing is not respected and I cant find a parameter that will help with that. The cloned voice sounds like the original but accelerated, where periods sound like commas and commas effectively vanish by the time you get close to the one minute mark.
- I lowered the input reference tempo to 20% slower than normal (using Audacity and others program), so it was pretty slow but the output was still the same speed and length, though the input speaking rate was much slower.
- I thought that maybe it was either related to the maximum new tokens, changing it from 2048 to 4096 but that didnt resolve the issue.
- I also though that maybe 1 minute of continuous text may have been too much, even with a token buffer of 4096, so I shortened the output to approximately 30 seconds but the speaking rate was still accelerated.
Reproduction
Utilize different reference text and have it speak a minute worth of text
Logs
Environment Information
Linux 22.04
Known Issue
Description
I am using a reference audio, cloned voice and the output for what should be approximately one minute of output is ALWAYS 41-48 seconds. I have tried a multitude of modifications based on the available parameters, including temperature, subtalker temperature, do_sample and subtalker_dosample. None of these parameters appear to have an effect on speaking rate.
Something that I did as a test was to change the input reference to 20% slower (using Audacity or other program), making the audio reference significantly slower than the original. Yet, the one minute of text outcome was still time compressed to 41-48 seconds and as you get closer to the end, pausing and stopping for punctuation is almost always greatly compressed.
The speaker pacing is not respected and I cant find a parameter that will help with that. The cloned voice sounds like the original but accelerated, where periods sound like commas and commas effectively vanish by the time you get close to the one minute mark.
Reproduction
Utilize different reference text and have it speak a minute worth of text
Logs
Environment Information
Linux 22.04
Known Issue