Releases: openvpi/TIFA
Releases · openvpi/TIFA
Release list
v1.0.0: pretrained model
This is TIFA 1.0 pretrained model.
Datasets
Voice datasets
- ~68h of private data with manual phoneme timing labels
- ~142h of public data with transcripts
- Children's Song Dataset (English subset)
- GTSinger (Chinese, English, and Japanese subsets, including paired speech)
- MIR-1K (vocal channel)
- OpenSinger
- Opencpop
- PopCS
- M4Singer
- No.7 Singing Database
Natural noise & accompaniments datasets
- DEMAND
- MUSAN (speech part excluded)
- MIR-1K
- MusicNet
- MUSDB18-HQ
- private Asian ethnic instruments dataset
Reverb dataset
Supported languages and special tags
| Language | Code | Accepted forms | Quality of coverage |
|---|---|---|---|
| Chinese Mandarin | zh | Hanzi, pinyin | Excellent |
| English | en | Words | Good |
| Japanese | ja | Kanji, kana, romaji | Good |
| Yue (Cantonese) | yue | Hanzi, jyutping | Moderate |
| Tag | Meaning |
|---|---|
| AP | Aspiration |
| EP | Expiration |
| GS | Glottal stop |
Specifications
- #Params: ~40M
- Max context length: ~1 minute (longer inputs may result in unstable alignment)
Acknowledgements
Data labels are provided by:
License
The model files apply CC BY-NC-SA 4.0 license. This license only constraints the model file itself, not the data that flows through it.