Skip to content
Discussion options

You must be logged in to vote

Thanks a lot for the detailed report!

What you're seeing is actually a known limitation of the current Silero VAD. In dense, continuous speech (such as interviews or conversations with short pauses), the model sometimes fails to separate adjacent speech segments and ends up merging them into a single long segment. Improving speech boundary detection in these scenarios is something we plan to address in future releases.

From our experience, FireRedVAD indeed tends to detect short pauses inside dense speech better than Silero. However, the trade-off is that it is noticeably less robust to background noise and other acoustic artifacts. Silero was intentionally tuned to be more conservative, …

Replies: 1 comment 1 reply

Comment options

You must be logged in to vote
1 reply
@chuducnam
Comment options

Answer selected by snakers4
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
help wanted Extra attention is needed
2 participants
Converted from issue

This discussion was converted from issue #785 on July 13, 2026 09:57.