Persistent Anti-Muslim Bias in Large Language Models
Abubakar Abid1, Maheen Farooqi2, James Zou3∗ 1Department of Electrical Engineering, Stanford University, CA, USA 2Department of Health Sciences, McMaster University, ON, Canada 3Department of Biomedical Data Science, Stanford University, CA, USA ∗To whom correspondence should be addressed:
Abstract
It has been observed that large-scale language models capture undesirable societal biases, e.g. relating to race and gender; yet religious bias has been relatively unexplored. We demonstrate that GPT-3, a state-of-the-art contextual language model, captures persistent Muslim-violence bias. We probe GPT-3 in various ways, including prompt completion, analogical reasoning, and story generation, to understand this anti-Muslim bias, demonstrating that it appears consistently and creatively in different uses of the model and that it is severe even compared to biases about other religious groups. For instance, “Muslim” is analogized to “terrorist” in 23% of test cases, while “Jewish” is mapped to “money” in 5% of test cases. We quantify the positive distraction needed to overcome this bias with adversarial text prompts, and find that use of the most positive 6 adjectives reduces violent completions for “Muslims” from 66% to 20%, but which is still higher than for other religious groups.
中文速览
大型语言模型(如GPT-3)会从海量互联网文本中悄悄吸收社会偏见,而针对伊斯兰教群体的偏见此前研究相对不足。研究者通过提示补全、类比推理和故事生成三种方式系统探测GPT-3,发现"穆斯林"这个词与暴力的关联极为显著且持续出现:在"两个穆斯林走进……"这一中性提示下,100次生成中有66次包含枪击、炸弹、谋杀等暴力内容;在类比测试中,"穆斯林"有23%的概率被类比为"恐怖分子",远高于其他宗教群体中任何单一刻板词汇的出现频率。研究者还尝试在提示中加入积极形容词来抵消偏见,效果最佳的6个词能将暴力生成比例从66%压低至20%,但这仍高于把"穆斯林"替换为其他宗教群体时的比例,且会产生将叙事引向特定方向的副作用。这项研究表明,强大语言模型所内化的偏见并非简单记忆,而是以创造性、多样化的方式渗透在各类应用场景中,这使得检测与消除偏见远比想象中困难,对依赖这类模型的实际产品具有重要警示意义。
原文 arXiv:2101.05783;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2101.05783v2