Auditing Black-box Models for Indirect Influence Thanks: This research was funded in part by the NSF under grants IIS-1251049, CNS-1302688, IIS-1513651, DMR-1307801, IIS-1633724, and IIS-1633387.
Philip Adler1, Casey Falk1, Sorelle A. Friedler1, Gabriel Rybeck1, Carlos Scheidegger2, Brandon Smith1, and Suresh Venkatasubramanian3 Affiliation: 1 Dept. of Computer Science, Haverford College, Haverford, PA, USA Affiliation: 2 Dept. of Computer Science, University of Arizona, Tucson, AZ, USA Affiliation: 3 Dept. of Computer Science, University of Utah, Salt Lake City, UT, USA
Abstract
Data-trained predictive models see widespread use, but for the most part they are used as black boxes which output a prediction or score. It is therefore hard to acquire a deeper understanding of model behavior, and in particular how different features influence the model prediction. This is important when interpreting the behavior of complex models, or asserting that certain problematic attributes (like race or gender) are not unduly influencing decisions.
原文 arXiv:1602.07043;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1602.07043v2