Speaker Diarization Model (pyannote segmentation-3.0)
·
170 commits
to master
since this release
pyannote segmentation-3.0 CoreML model for speaker diarization. Input: 10s mono 16kHz audio [1,1,160000]. Output: [1,589,7] speaker activity logits (powerset encoding: 3 speakers + overlaps). 5.8MB. License: MIT.