-
Notifications
You must be signed in to change notification settings - Fork 76
Available data sources
On this page we list the scripts for data sources. Some of these are fully functional, others are deprecated. Let us know if you have a new data source to add.
For datasource-specific information, check the README files in the folder of the respective data source.
|----------------------|------------------------|----------------------|-------|---------------------------------------------------------------------------------------------------------------------------------------------| | 4chan | 4chan API | Comments + OPs | Yes | We wrote several scripts to import data from 4chan archives in the helper-scripts folder, e.g. this script to import csv dumps from 4plebs. | | 8chan | 4chan API | Comments + OPs | Yes | 8chan is now defunct. We scraped live data when it was still online. Let us know in case you are interested in a database copy. | | 8kun | 8chan API | Comments + OPs | Yes | Similar to the 4chan data source. | | Bitchute | Scraping | Videos + comments | No | Uses BitChutes web search endpoint, and scrapes data from the live website. | | Douban | Scraping | Comments + OPs | No | Small datasets can be collected; due to rate-limiting, large searches may not complete properly. | | Import from tool | Files from other tools | - | No | This to import files from tools like CrowdTangle. | | Instagram | ZeeSchuimer | Posts | No | Must be actively scraped via your browser and the Zeeschuimer plugin. | | LinkedIn | ZeeSchuimer | Posts | No | Must be actively scraped via your browser and the Zeeschuimer plugin. | | Parler | Parler API | Posts | No | Uses Parler's unofficial web API; requires a valid Parler login for usage. | | Parliament | PENELOPE API | Speeches | No | Static snapshot of transcripts of parliamentary speeches from the UK and Germany. | | Reddit | Pushshift API | Comments + OPs | No | Data retrieved via Pushshift. | | Telegram | Telegram API | Messages in open groups | No | Requires a personal API key, which can be obtained by anyone with a Telegram account here. | | The Guardian climate change | - | Articles + comments | Yes | A static data source of The Guardian articles and comments concerning climate change (collected by VUB). | | TikTok | Zeeschuimer | Posts | No | Must be actively scraped via your browser and the Zeeschuimer plugin. | | Tumblr | Tumblr API | Posts + text reblogs | No | Requires API keys which you can obtain here | | Twitter | Twitter v2 API | Tweets | No | Requires an API key for either the standard or academic search API. The standard API is limited to the most recent 7 days of tweets. For keys apply here. | | Usenet | - | Comments + OPs | Yes | Requires a local, static Usenet database. |