Towards A Rigorous Science of Interpretable Machine Learning
Finale Doshi-Velez Been Kim Note: Authors contributed equally.
Abstract
From autonomous cars and adaptive email-filters to predictive policing systems, machine learning (ML) systems are increasingly ubiquitous; they outperform humans on specific tasks (Mnih et al. 2013; Silver et al. 2016; Hamill 2017) and often guide processes of human understanding and decisions (Carton et al. 2016; Doshi-Velez et al. 2014). The deployment of ML systems in complex applications has led to a surge of interest in systems optimized not only for expected task performance but also other important criteria such as safety (Otte 2013; Amodei et al. 2016; Varshney and Alemzadeh 2016), nondiscrimination (Bostrom and Yudkowsky 2014; Ruggieri et al. 2010; Hardt et al. 2016), avoiding technical debt (Sculley et al. 2015), or providing the right to explanation (Goodman and Flaxman 2016). For ML systems to be used safely, satisfying these auxiliary criteria is critical. However, unlike measures of performance such as accuracy, these criteria often cannot be completely quantified. For example, we might not be able to enumerate all unit tests required for the safe operation of a semi-autonomous car or all confounds that might cause a credit scoring system to be discriminatory. In such
原文 arXiv:1702.08608;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1702.08608v2