A new collection of codes to elaborate the dataset named "QASports 2.0", the first large sports question answering dataset for open questions. QASports 2.0 contains real data of players, teams and matches from the top sports around the globe (soccer, football, basket, rugby, cricket, etc). It counts million questions and answers, cleaned and organized documents from Wikipedia-like sources. It contains:
- Crawler for Fandom wiki pages
- Fetching the list of useful links in a Fandom wiki
- Processing techniques to clean and transform the text
- Question-answering context extracting script
- Question-answering automatic dataset generation
- Data section algorithms for representative questions