Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity AnnotationConference: Taipei ’22; CSCW 2022;CCS: Human-centered computing Empirical studies in HCI
Nitesh Goyal email: OrcID: 0000-0002-4666-1926 Affiliation: Google Research, Google , 111 8th Ave , New York , NY , USA , 11201 , Ian D. Kivlichan email: OrcID: 0000-0003-2719-2500 Affiliation: Jigsaw, Google , 111 8th Ave , New York , NY , USA , 11201 , Rachel Rosen email: OrcID: 0000-0003-2927-1245 Affiliation: Jigsaw, Google , 111 8th Ave , New York , NY , USA , 11201 and Lucy Vasserman OrcID: 0000-0002-6938-0713 email: Affiliation: Jigsaw, Google , 111 8th Ave , New York , NY , USA , 11201
Abstract
Machine learning models are commonly used to detect toxicity in online conversations. These models are trained on datasets annotated by human raters. We explore how raters’ self-described identities impact how they annotate toxicity in online comments. We first define the concept of specialized rater pools: rater pools formed based on raters’ self-described identities, rather than at random. We formed three such rater pools for this study–specialized rater pools of raters from the U.S. who identify as African American, LGBTQ, and those who identify as neither. Each of these rater pools annotated the same set of comments, which contains many references to these identity groups. We found that rater identity is a statistically significant factor in how raters will annotate toxicity for identity-related annotations. Using preliminary content analysis, we examined the comments with the most disagreement between rater pools and found nuanced differences in the toxicity annotations. Next, we trained models on the annotations from each of the different rater pools, and compared the scores of these models on comments from several test sets. Finally, we discuss how using raters that self-ide
原文 arXiv:2205.00501;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.00501v1