Handling Bias in Toxic Speech Detection: A Survey
Tanmay Garg IIIT DelhiIndia , Sarah Masud IIIT DelhiIndia , Tharun Suresh IIIT DelhiIndia and Tanmoy Chakraborty IIT DelhiIndia
Abstract
Detecting online toxicity has always been a challenge due to its inherent subjectivity. Factors such as the context, geography, socio-political climate, and background of the producers and consumers of the posts play a crucial role in determining if the content can be flagged as toxic. Adoption of automated toxicity detection models in production can thus lead to a sidelining of the various groups they aim to help in the first place. It has piqued researchers’ interest in examining unintended biases and their mitigation. Due to the nascent and multi-faceted nature of the work, complete literature is chaotic in its terminologies, techniques, and findings. In this paper, we put together a systematic study of the limitations and challenges of existing methods for mitigating bias in toxicity detection.
中文速览
网络有毒言论(toxic speech)的自动检测面临一个棘手的问题:模型往往会对特定种族、性别或身份群体产生"非预期偏见",反而伤害了它本应保护的弱势群体。这篇综述系统梳理了近年来研究者在评估和缓解这类偏见方面提出的方法,按照"伤害来源"(采样偏差、词汇偏差、标注偏差)和"伤害对象"(种族、性别、政治心理偏见)两条主线构建了一套分类框架,并通过一个案例研究揭示了"偏见迁移"(bias shift)现象——基于知识的去偏方法在消除某一偏见的同时可能将偏见转移到其他群体身上。研究发现,现有方法效果普遍有限,根本原因在于有毒言论检测本质上是一项高度主观的任务,与文本蕴含等客观NLP任务有本质差异。这一综述有助于研究者厘清该领域混乱的术语与技术路线,为构建更公平、更鲁棒的内容审核系统指明方向。
原文 arXiv:2202.00126;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2202.00126v3