The Cross-Lingual Sentiment (CLS) dataset comprises about 800.000 Amazon product reviews in the four languages English, German, French, and Japanese.
For more information on the construction of the dataset see (Prettenhofer and Stein, 2010) or the enclosed readme files. If you have a question after reading the paper and the readme files, please contact Peter Prettenhofer.
We provide the dataset in two formats: 1) a processed format which corresponds to the preprocessing (tokenization, etc.) in (Prettenhofer and Stein, 2010); 2) an unprocessed format which contains the full text of the reviews (e.g., for machine translation or feature engineering).
To download the corpus use the following link:
(230 MB, MD5 sum: 3d5d7b4fe3945e0f5cf438dd15e1c25d). [readme]
(300 MB, MD5 sum: 3956bae48add21162ffefb26ae19b266). [readme]
If you use the dataset in your research, please send us a copy of your publication. We kindly ask you to refer to the corpus via [bib].
The dataset was first used by (Prettenhofer and Stein, 2010). It consists of Amazon product reviews for three product categories---books, dvds and music---written in four different languages: English, German, French, and Japanese. The German, French, and Japanese reviews were crawled from Amazon in November, 2009. The English reviews were sampled from the Multi-Domain Sentiment Dataset (Blitzer et. al., 2007). For each language-category pair there exist three sets of training documents, test documents, and unlabeled documents. The training and test sets comprise 2.000 documents each, whereas the number of unlabeled documents varies from 9.000 - 170.000.
For more information on the construction of the dataset see (Prettenhofer and Stein, 2010) and the enclosed readme files.
We kindly thank Mark Dredze and John Blitzer for the permission to include a sample of the Multi-Domain Sentiment Dataset (Blitzer et. al., 2007) in our dataset.