Please use this identifier to cite or link to this item: https://hdl.handle.net/2445/220216
Title: NewsCom-TOX: A corpus of comments on news articles annotated for toxicity in Spanish
Author: Taulé Delor, Mariona
Nofre, Montserrat
Bargiela, Víctor
Bonet, Xavier
Keywords: Telenotícies
Fake news
Castellà (Llengua)
Corpus (Lingüística)
Television broadcasting of news
Fake news
Spanish language
Corpora (Linguistics)
Issue Date: 17-Jan-2024
Publisher: Springer Verlag
Abstract: In this article, we present the NewsCom-TOX corpus, a new corpus manually annotated for toxicity in Spanish. NewsCom-TOX consists of 4359 comments in Spanish posted in response to 21 news articles on social media related to immigration, in order to analyse and identify messages with racial and xenophobic content. This corpus is multi-level annotated with different binary linguistic categories -stance, target, stereotype, sarcasm, mockery, insult, improper language, aggressiveness and intolerance- taking into account not only the information conveyed in each comment, but also the whole discourse thread in which the comment occurs, as well as the information conveyed in the news article, including their images. These categories allow us to identify the presence of toxicity and its intensity, that is, the level of toxicity of each comment. All this information is available for research purposes upon request. Here we describe the NewsCom-TOX corpus, the annotation tagset used, the criteria applied and the annotation process carried out, including the inter-annotator agreement tests conducted. A quantitative analysis of the results obtained is also provided. NewsCom-TOX is a linguistic resource that will be valuable for both linguistic and computational research in Spanish in NLP tasks for the detection of toxic information.
Note: Versió postprint del document publicat a: https://doi.org/10.1007/s10579-023-09711-x
It is part of: Language Resources And Evaluation, 2023, num.58, p. 1115-1155
URI: https://hdl.handle.net/2445/220216
Related resource: https://doi.org/10.1007/s10579-023-09711-x
ISSN: 1574-020X
Appears in Collections:Articles publicats en revistes (Filologia Catalana i Lingüística General)

Files in This Item:
File Description SizeFormat 
840381.pdf993.46 kBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.