Repository navigation
Is there a speaker recognition gap? #730
Replies: 2 comments 1 reply
|
We already have reusable speaker encoders in the framework, including ECAPA-TDNN, TitaNet, and CAM++, but they aren’t currently exposed as standalone speaker-recognition models through the server. So there is a gap in the user-facing support. Let me think about the best way to do this. Have you tested other models? Moving inference to the server could help thin-client latency, but improving recognition on short utterances would need separate evaluation. |
|
Grok tells me that there are a few models I should consider:
I have not figured out how to get any of these working yet. I don't think there's any way such models can be supported without the ability to make speaker profiles. Personally I only need one profile, but it would be particularly cool one day if an assistant could automatically generate them when it encounters a new voice, and then recognize that person in the future. |
Uh oh!
There was an error while loading. Please reload this page.
I could be missing something, but it appears there are no easy, established ways to do speaker recognition. (Speaker recognition is comparing a voice sample to an enrolled profile to match a specific individual.)
I currently do this with ERes2Net sherpa-onnx. It's very bad with short utterances and introduces latency on thin clients. Running this in audio.cpp on the server would be vastly better. This would also make it possible to run bigger models.
All reactions