Cyberbullying Detection with Fairness Constraints
Oguzhan Gencoglu Tampere University, Faculty of Medicine and Health Technology, Tampere, Finland
Abstract
Cyberbullying is a widespread adverse phenomenon among online social interactions in today’s digital society. While numerous computational studies focus on enhancing the cyberbullying detection performance of machine learning algorithms, proposed models tend to carry and reinforce unintended social biases. In this study, we try to answer the research question of “Can we mitigate the unintended bias of cyberbullying detection models by guiding the model training with fairness constraints?”. For this purpose, we propose a model training scheme that can employ fairness constraints and validate our approach with different datasets. We demonstrate that various types of unintended biases can be successfully mitigated without impairing the model quality. We believe our work contributes to the pursuit of unbiased, transparent, and ethical machine learning solutions for cyber-social health.
中文速览
网络欺凌(cyberbullying)自动检测模型在训练时会悄悄"继承"并放大社会偏见——例如把含有"gay""jew"等词的无害言论错误判定为仇恨内容,或对非裔英语推文给出更严苛的评分。本研究将公平性评估指标(假阳性/假阴性群体差异)直接转化为约束条件,嵌入神经网络的训练优化过程,让模型在学习识别欺凌内容的同时被强制引导向各群体表现更均等的方向收敛。研究者在四个来源各异的数据集上验证了该方案,涵盖性别偏见、语言偏见、时间偏见和宗教/种族/国籍偏见等不同场景,结果显示各类偏见均得到有效缓解,而整体检测性能不仅未下降,部分指标反有提升。这项工作的意义在于提供了一种无需修改训练数据、可与现有去偏方法叠加使用、且推断时不依赖群体标签的通用公平训练框架,为构建更透明、更负责任的内容安全系统提供了实践路径。
原文 arXiv:2005.06625;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2005.06625v2