Evaluating and Mitigating Discrimination in Language Model Decisions
Alex Tamkin Affiliation: Anthropic, San Francisco, USA Correspondence to: Amanda Askell Affiliation: Anthropic, San Francisco, USA Liane Lovitt Affiliation: Anthropic, San Francisco, USA Esin Durmus Affiliation: Anthropic, San Francisco, USA Nicholas Joseph Affiliation: Anthropic, San Francisco, USA Shauna Kravec Affiliation: Anthropic, San Francisco, USA Karina Nguyen Affiliation: Anthropic, San Francisco, USA Jared Kaplan Affiliation: Anthropic, San Francisco, USA Deep Ganguli Affiliation: Anthropic, San Francisco, USA
Abstract
As language models (LMs) advance, interest is growing in applying them to high-stakes societal decisions, such as determining financing or housing eligibility. However, their potential for discrimination in such contexts raises ethical concerns, motivating the need for better methods to evaluate these risks. We present a method for proactively evaluating the potential discriminatory impact of LMs in a wide range of use cases, including hypothetical use cases where they have not yet been deployed. Specifically, we use an LM to generate a wide array of potential prompts that decision-makers may input into an LM, spanning 70 diverse decision scenarios across society, and systematically vary the demographic information in each prompt. Applying this methodology reveals patterns of both positive and negative discrimination in the Claude 2.0 model in select settings when no interventions are applied. While we do not endorse or permit the use of language models to make automated decisions for the high-risk use cases we study, we demonstrate techniques to significantly decrease both positive and negative discrimination through careful prompt engineering, providing pathways toward safer depl
原文 arXiv:2312.03689;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2312.03689v1