A Survey on Bias and Fairness in Machine Learning
Ninareh Mehrabi , Fred Morstatter , Nripsuta Saxena , Kristina Lerman and Aram Galstyan, USC-ISI
Abstract
With the widespread use of artificial intelligence (AI) systems and applications in our everyday lives, accounting for fairness has gained significant importance in designing and engineering of such systems. AI systems can be used in many sensitive environments to make important and life-changing decisions; thus, it is crucial to ensure that these decisions do not reflect discriminatory behavior toward certain groups or populations. More recently some work has been developed in traditional machine learning and deep learning that address such challenges in different subdomains. With the commercialization of these systems, researchers are becoming more aware of the biases that these applications can contain and are attempting to address them. In this survey we investigated different real-world applications that have shown biases in various ways, and we listed different sources of biases that can affect AI applications. We then created a taxonomy for fairness definitions that machine learning researchers have defined in order to avoid the existing bias in AI systems. In addition to that, we examined different domains and subdomains in AI showing what researchers have observed with reg
中文速览
机器学习算法如今已被广泛用于招聘、贷款、司法量刑等高风险决策场景,但这些系统常常对特定群体表现出系统性歧视——比如风险评估工具COMPAS对非裔美国人误判率偏高、人脸识别系统对深肤色人群准确率更低。这篇综述从数据、算法、用户交互三个层面梳理了AI偏见(bias)的来源:训练数据本身可能携带测量偏差、代表性不足等问题,算法设计本身也会放大这些偏差,而有偏的输出反过来又影响用户行为,形成恶性反馈循环。作者进一步整理了机器学习领域研究者提出的各类公平性(fairness)定义分类体系,并梳理了自然语言处理、计算机视觉、推荐系统等多个子领域中检测和缓解偏见的最新方法。这项工作为希望在各自领域落地公平AI的研究者提供了系统性的参考路线图,也指出了仍有大量亟待解决的开放问题。
原文 arXiv:1908.09635;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1908.09635v3