Skip to content

EEND max_num_speakers of ALL audio

Latest

Choose a tag to compare

@marccasals98 marccasals98 released this 28 Mar 11:18
· 7 commits to main since this release
6d4b022

In this release, we make the first implementation of EEND system. The data.py has been managed the following way: Given an audio $X$, we create the label $Y$ with $\textbf{y}_t = [ y_{t,c} \in \{ 0,1\} \mid t = 1,\dots, C]$. In this release, we consider $C$ speakers as all the speakers comprehended in the audio X.

This can lead to the following problem: Imagine we have an audio in which 20 speakers speak. However, in each segment, only two speakers speak within it. Then the labels are full of zeros since $C=20$. The future approach will consider individual $C$ for each segment, so in this case $C=2$.