Towards A Rigorous Science of Interpretable Machine Learning
Finale Doshi-Velez∗ and Been Kim111Authors contributed equally.
Abstract
From autonomous cars and adaptive email-filters to predictive policing systems, machine learning (ML) systems are increasingly ubiquitous; they outperform humans on specific tasks (Mnih et al., 2013; Silver et al., 2016; Hamill, 2017) and often guide processes of human understanding and decisions (Carton et al., 2016; Doshi-Velez et al., 2014). The deployment of ML systems in complex applications has led to a surge of interest in systems optimized not only for expected task performance but also other important criteria such as safety (Otte, 2013; Amodei et al., 2016; Varshney and Alemzadeh, 2016), nondiscrimination (Bostrom and Yudkowsky, 2014; Ruggieri et al., 2010; Hardt et al., 2016), avoiding technical debt (Sculley et al., 2015), or providing the right to explanation (Goodman and Flaxman, 2016). For ML systems to be used safely, satisfying these auxiliary criteria is critical. However, unlike measures of performance such as accuracy, these criteria often cannot be completely quantified. For example, we might not be able to enumerate all unit tests required for the safe operation of a semi-autonomous car or all confounds that might cause a credit scoring system to be discrimina
中文速览
机器学习系统在自动驾驶、医疗决策等高风险场景中越来越普遍,但"可解释性(interpretability)"究竟是什么、该怎么评估,学界至今没有共识。这篇论文从根本上追问为何需要可解释性——核心论点是:当问题定义本身存在"不完整性(incompleteness)",即无法用完整的量化目标覆盖安全、公平、因果等关切时,可解释性就成为弥补这一缺口的关键手段。作者提出了一套三层评估分类框架:以真实任务和领域专家为核心的"应用导向评估"、以简化任务和普通用户为核心的"人本评估"、以及无需人类参与的"功能代理评估",并明确指出三者的适用条件与权衡。这项工作的重要性在于,它为快速扩张的可解释性研究提供了概念地基和方法论规范,有助于研究者在不同场景下选择合适的评估方式,推动该领域从"感觉说得通"走向可比较、可复现的严格科学。
原文 arXiv:1702.08608;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1702.08608v2