The SOMTUME dataset contains textual information gathered from social media and news sites, comprising two segments: Trustworthiness Information Content (TIC) and Uncertain Information Content (UIC). The texts pertain to the migration of Ukrainians to the European Union from February 22, 2022, to the present day. The TIC section encompasses texts and recognized audio files sourced from ten government digital platforms in Ukraine, Poland, Canada, and the USA, amounting to 510 publications with a total size of 630,188 tokens. The UIC section comprises texts from the Telegram social network, involving 1,578,020 short text messages extracted from 103 chats across 40 countries. The SOMTUME dataset includes columns such as 'Country,' 'Date,' 'Language,' 'Lemmatized Text,' and 'Link'. However, to comply with copyright and ensure private data protection, the dataset doesn't include the full texts of articles or direct links to the Telegram groups.
Distributions of the TIC part of dataset texts by languages and sources are linked below.
| Sources link | Percentage of texts |
|---|---|
| https://dtm.iom.int/reports | 25% |
| https://www.gov.pl/ | 14% |
| https://dmsu.gov.ua/ | 13% |
| https://www.youtube.com/ | 10% |
| https://twitter.com/ | 4% |
| https://www.uscis.gov/ | 2% |
| https://ukraine.iom.int/uk | < 1% |
| https://ccrweb.ca/en | < 1% |
| Language | Percentage |
|---|---|
| English | 68% |
| Ukrainian | 14% |
| Polish | 7% |
MIT
Free Software, Hell Yeah!