European Multilingual News Articles Dataset with Topic Annotation

DOI10.5281/zenodo.10397400Zenodo10397400MaRDI QIDQ6709015FDOQ6709015

Dataset published at Zenodo repository.

Lorenzo Bellomo, Virginia Morini, Giulio Rossetti, Dino Pedreschi, Paolo Ferragina

Publication date: 17 December 2023

Copyright license: Creative Commons Attribution 4.0 International

The European Multilingual News Articles Dataset is composed of over 18 million European news articles coming from 205 media outlets belonging to 27 European countries (i.e., all EU countries belonging to the European Union) with the addition of the United Kingdom. Articles range in a time period from 2017 to 2021 and are written in their original languages, for a total of 23 different languages included. After selecting reliable, nationwide European media outlets, each article (i.e., title, textual content, URL, and date and time of publication) was extracted from the Common Crawl News Corpus, which contains petabytes of raw web page data collected since 2016. The dataset is released without any text pre-processing other than a cleanup of XML tags. Further, we enriched it by adding several media metadata (e.g., frequency of publication, distribution area, language, type of media). Moreover, we enhanced the dataset by adding - whenever possible - article-level topic annotation by using articles' URLs as a proxy of the topic discussed. In the end, we were able to assign a topic to over 4 million articles (33 unique topics, e.g., politics, sport, entertainment), thus 23.2% of the entire dataset. Further, from URLs, we also extract the types of over 4 million articles (15 unique article types, e.g., news, international, multimedia).

This page was built for dataset: European Multilingual News Articles Dataset with Topic Annotation