The Tail Wagging the Dog: Dataset Construction Biases of Social Bias Benchmarks
Nikil Roashan Selvam11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Sunipa Dev22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Daniel Khashabi33{}^{3}start_FLOATSUPERSCRIPT 3 end_FLOATSUPERSCRIPT Tushar Khot44{}^{4}start_FLOATSUPERSCRIPT 4 end_FLOATSUPERSCRIPT Kai-Wei Chang11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT 11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPTUniversity of California, Los Angeles 22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPTGoogle Research 33{}^{3}start_FLOATSUPERSCRIPT 3 end_FLOATSUPERSCRIPTJohns Hopkins University 44{}^{4}start_FLOATSUPERSCRIPT 4 end_FLOATSUPERSCRIPTAllen Institute for AI
Abstract
How reliably can we trust the scores obtained from social bias benchmarks as faithful indicators of problematic social biases in a given model? In this work, we study this question by contrasting social biases with non-social biases that stem from choices made during dataset construction (which might not even be discernible to the human eye). To do so, we empirically simulate various alternative constructions for a given benchmark based on seemingly innocuous modifications (such as paraphrasing or random-sampling) that maintain the essence of their social bias. On two well-known social bias benchmarks (Winogender and BiasNLI), we observe that these shallow modifications have a surprising effect on the resulting degree of bias across various models and consequently the relative ordering of these models when ranked by measured bias. We hope these troubling observations motivate more robust measures of social biases.
中文速览
衡量语言模型社会偏见(social bias)的基准测试,其测量结果究竟有多可靠?研究者选取 Winogender 和 BiasNLI 这两个知名基准,通过在保留社会偏见内涵的前提下对数据集做看似无关紧要的修改——例如对动词取反、替换同义词、添加形容词或子句、以及随机重新采样——来构造多种替代版本,从而考察这些表面改动对偏见测量结果的影响。实验发现,这些细微改动会使同一模型测出的偏见分数出现大幅波动,例如仅对 BiasNLI 中的动词添加否定,ELMo-DA 模型的偏见指标就从 41.64 骤降至 13.40,甚至导致不同模型在"谁更有偏见"这一排名上发生逆转。这说明现有社会偏见基准在很大程度上也在无意间测量着数据集构造方式引入的非社会性偏见,两者难以分离,因此用这些基准对模型进行横向比较或部署决策时需要格外谨慎,也亟需开发更加稳健的社会偏见评估方法。
原文 arXiv:2210.10040;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2210.10040v2