Counterfactual Instances Explain Little
Adam White, Artur d’Avila Garcez City Data Science Institute City, University of London, London, EC1V 0HB, UK {Adam.White, A.Garcez
Abstract
In many applications, it is important to be able to explain the decisions of machine learning systems. An increasingly popular approach has been to seek to provide counterfactual instance explanations. These specify close possible worlds in which, contrary to the facts, a person receives their desired decision from the machine learning system. This paper will draw on literature from the philosophy of science to argue that a satisfactory explanation must consist of both counterfactual instances and a causal equation (or system of equations) that support the counterfactual instances. We will show that counterfactual instances by themselves explain little. We will further illustrate how explainable AI methods that provide both causal equations and counterfactual instances can successfully explain machine learning predictions.
中文速览
机器学习系统的"反事实实例解释"(counterfactual instance explanation)——即告诉用户"如果你的收入再高3000美元,贷款就会获批"——是当前可解释AI领域的主流方法,但这类解释究竟够不够用,一直缺乏严格的哲学审视。本文借鉴科学哲学中以Woodward为代表的因果解释理论,指出单靠反事实实例根本无法构成令人满意的解释,因为它既不揭示变量之间的因果结构,也无法保证给出的行动建议在现实中真正可行。作者主张,一个完整的解释必须同时包含两部分:支持反事实推断的不变因果方程(invariant causal equation),以及由该方程所推导出的反事实实例。文章进一步展示,那些既提供因果方程又给出反事实实例的XAI方法,能够真正满足科学解释的标准,从而帮助用户理解预测背后的因果机制、并获得切实可行的改变路径。这一结论对整个可解释AI领域具有重要意义,提示研究者不能止步于生成"近邻反事实点",而应将因果建模纳入解释框架的核心。
原文 arXiv:2109.09809;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2109.09809v1