On Measures of Biases and Harms in NLP
Sunipa Dev1 Emily Sheng2∗ Jieyu Zhao1∗ Aubrie Amstutz∗ \ANDJiao Sun3 Yu Hou3 Mattie Sanseverino1 Jiin Kim1 Akihiro Nishi1\ANDNanyun Peng1,3 Kai-Wei Chang1 1University of California, Los Angeles, 2Microsoft Research, 3University of Southern California equal contribution
Abstract
Recent studies show that Natural Language Processing (NLP) technologies propagate societal biases about demographic groups associated with attributes such as gender, race, and nationality. To create interventions and mitigate these biases and associated harms, it is vital to be able to detect and measure such biases. While existing works propose bias evaluation and mitigation methods for various tasks, there remains a need to cohesively understand the biases and the specific harms they measure, and how different measures compare with each other. To address this gap, this work presents a practical framework of harms and a series of questions that practitioners can answer to guide the development of bias measures. As a validation of our framework and documentation questions, we also present several case studies of how existing bias measures in NLP—both intrinsic measures of bias in representations and extrinsic measures of bias of downstream applications—can be aligned with different harms and how our proposed documentation questions facilitates more holistic understanding of what bias measures are measuring.
中文速览
NLP系统会把社会中对性别、种族、国籍等群体的偏见"学进去"并放大,但现有的偏见度量方法五花八门,研究者往往说不清楚自己到底在测量哪种伤害、测量结果之间又如何比较。这篇文章提出了一套实用框架,将NLP偏见度量与五类具体伤害——刻板印象(Stereotyping)、贬损(Disparagement)、非人化(Dehumanization)、抹除(Erasure)和服务质量不公平(Quality of Service)——系统对应起来,并给出一组"文档化问题",帮助研究者在设计或使用偏见度量时厘清所针对的群体、数据集局限和指标背后的假设。作者用43个已有偏见度量(涵盖词向量内在测量和下游任务外在测量)做了标注和案例分析,展示了同一测量工具可能无意中混淆多种伤害、而看似相同的任务定义实则指向不同伤害的现象。这项工作的价值在于为整个领域提供了一套共同语言,让不同偏见度量之间的比较和干预效果的评估都变得更加严格和可操作。
原文 arXiv:2108.03362;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2108.03362v2