The Tail Wagging the Dog: Dataset Construction Biases of Social Bias Benchmarks
Nikil Roashan Selvam11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Sunipa Dev22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Daniel Khashabi33{}^{3}start_FLOATSUPERSCRIPT 3 end_FLOATSUPERSCRIPT Tushar Khot44{}^{4}start_FLOATSUPERSCRIPT 4 end_FLOATSUPERSCRIPT Kai-Wei Chang11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT 11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPTUniversity of California, Los Angeles 22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPTGoogle Research 33{}^{3}start_FLOATSUPERSCRIPT 3 end_FLOATSUPERSCRIPTJohns Hopkins University 44{}^{4}start_FLOATSUPERSCRIPT 4 end_FLOATSUPERSCRIPTAllen Institute for AI
Abstract
How reliably can we trust the scores obtained from social bias benchmarks as faithful indicators of problematic social biases in a given model? In this work, we study this question by contrasting social biases with non-social biases that stem from choices made during dataset construction (which might not even be discernible to the human eye). To do so, we empirically simulate various alternative constructions for a given benchmark based on seemingly innocuous modifications (such as paraphrasing or random-sampling) that maintain the essence of their social bias. On two well-known social bias benchmarks (Winogender and BiasNLI), we observe that these shallow modifications have a surprising effect on the resulting degree of bias across various models and consequently the relative ordering of these models when ranked by measured bias. We hope these troubling observations motivate more robust measures of social biases.
原文 arXiv:2210.10040;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2210.10040v2