https://elevenlabs.io/docs/api-reference/text-to-speech/convert-with-timestamps their API allows returning letter-level timing. This would be a great improvement to have by default so we can have word-level timestamps, such that interrupts capture the parts of the words that were spoken.
https://elevenlabs.io/docs/api-reference/text-to-speech/convert-with-timestamps their API allows returning letter-level timing. This would be a great improvement to have by default so we can have word-level timestamps, such that interrupts capture the parts of the words that were spoken.