The Webis-Sentences-17 corpus is a collection of 3,369,618,811 sentences extracted from the ClueWeb12 web crawl. It is designed to allow for statistical analyses of human-written sentences. More details on the sentence extraction can be found in the associated publication.
The Webis-Simple-Sentences-17 corpus contains 471,085,690 English sentences from the Webis-Sentences-17 corpus. The sentences were sampled to achieve a level of sentence complexity similar to the one of sentences that humans make up as a memory aid for remembering passwords. Sentence complexity was determined by syllables per word.
Both corpora are split in training and test set as they are used in the associated publication. The test set is extracted from part 00 of the ClueWeb12, while the training set is extracted from the other parts.
You can access the Webis-Simple-Sentences-17 corpus on Zenodo.
You can request access to the Webis-Sentences-17 corpus (nearly 200 GB) by sending a mail to firstname.lastname@example.org.
If you use the dataset in your research, please send us a copy of your publication. We kindly ask you to refer to the corpus via [bib].