BBQ: A Hand-Built Bias Benchmark for Question Answering
Alicia Parrish Angelica Chen Affiliation: New York UniversityDept. of Linguistics Nikita Nangia Affiliation: New York UniversityCenter for Data Science Vishakh Padmakumar Affiliation: New York UniversityCenter for Data Science Affiliation: New York UniversityCenter for Data Science Jason Phang Jana Thompson Affiliation: New York UniversityCenter for Data Science Phu Mon Htut Affiliation: New York UniversityCenter for Data Science Samuel R. Bowman Affiliation: New York UniversityDept. of Linguistics Affiliation: New York UniversityCenter for Data Science Affiliation: New York UniversityCenter for Data Science Affiliation: New York UniversityDept. of Computer ScienceCorrespondence: {alicia.v.parrish,
Abstract
It is well documented that NLP models learn social biases, but little work has been done on how these biases manifest in model outputs for applied tasks like question answering (QA). We introduce the Bias Benchmark for QA (BBQ), a dataset of question sets constructed by the authors that highlight attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts. Our task evaluates model responses at two levels: (i) given an under-informative context, we test how strongly responses reflect social biases, and (ii) given an adequately informative context, we test whether the model’s biases override a correct answer choice. We find that models often rely on stereotypes when the context is under-informative, meaning the model’s outputs consistently reproduce harmful biases in this setting. Though models are more accurate when the context provides an informative answer, they still rely on stereotypes and average up to 3.4 percentage points higher accuracy when the correct answer aligns with a social bias than when it conflicts, with this difference widening to over 5 points on examples targeting gender for most m
原文 arXiv:2110.08193;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2110.08193v2