A monolingual parallel corpus for sentence simplification
Branch: master
Clone or download
Fetching latest commit…
Cannot retrieve the latest commit at this time.
Permalink
Type Name Latest commit message Commit time
Failed to load latest commit information.
README.md
sscorpus.gz

README.md

sscorpus: A monolingual parallel corpus for sentence simplification

This corpus contains 492,993 aligned sentences extracted by pairing Simple English Wikipedia with English Wikipedia. These source data were downloaded in May 2016.

The form of each line in the corpus: original sentence <TAB> simple sentence <TAB> similarity score

For questions, please contact Tomoyuki Kajiwara at Tokyo Metropolitan University.