“I’m sorry to hear that”: Finding New Biases in Language Models with a Holistic Descriptor Dataset
Eric Michael Smith Melissa Hall Melanie Kambadur Eleonora Presani Adina Williams Meta AI
Abstract
As language models grow in popularity, it becomes increasingly important to clearly measure all possible markers of demographic identity in order to avoid perpetuating existing societal harms. Many datasets for measuring bias currently exist, but they are restricted in their coverage of demographic axes and are commonly used with preset bias tests that presuppose which types of biases models can exhibit. In this work, we present a new, more inclusive bias measurement dataset, HolisticBias, which includes nearly 600 descriptor terms across 13 different demographic axes. HolisticBias was assembled in a participatory process including experts and community members with lived experience of these terms. These descriptors combine with a set of bias measurement templates to produce over 450,000 unique sentence prompts, which we use to explore, identify, and reduce novel forms of bias in several generative models. We demonstrate that HolisticBias is effective at measuring previously undetectable biases in token likelihoods from language models, as well as in an offensiveness classifier. We will invite additions and amendments to the dataset, which we hope will serve as a basis for more eas
中文速览
语言模型在生成文本时会对不同人群表现出系统性偏见,但现有的偏见评测数据集覆盖的人群类别太窄、评测方式也过于固定,导致很多针对边缘群体的偏见根本检测不到。为此,研究者发布了HolisticBias数据集,通过邀请来自不同社群的专家和亲历者共同参与,整理出横跨13个人口属性维度(如种族、性别、残障状况等)的近600个描述词,并将其与26种句子模板组合,生成逾45万条独特提示句。用这些提示句测试GPT-2、RoBERTa、DialoGPT和BlenderBot 2.0等模型后,研究者发现了此前被忽视的新型偏见——例如模型对残障人士的表述会触发过度同情语气,并验证该数据集能有效改善毒性分类器对不同群体的差异化判定问题。HolisticBias以开源"活文档"形式持续更新,为NLP领域提供了一套更全面、更标准化的社会偏见评测基础设施。
原文 arXiv:2205.09209;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.09209v2