ML Interpretability: Simple Isn’t Easy
Tim Räz111University of Bern, Institute of Philosophy, Länggassstrasse 49a, 3012 Bern, Switzerland. E-mail:
Abstract
The interpretability of ML models is important, but it is not clear what it amounts to. So far, most philosophers have discussed the lack of interpretability of black-box models such as neural networks, and methods such as explainable AI that aim to make these models more transparent. The goal of this paper is to clarify the nature of interpretability by focussing on the other end of the “interpretability spectrum”. The reasons why some models, linear models and decision trees, are highly interpretable will be examined, and also how more general models, MARS and GAM, retain some degree of interpretability. I find that while there is heterogeneity in how we gain interpretability, what interpretability is in particular cases can be explicated in a clear manner.
中文速览
机器学习模型的"可解释性"(interpretability)究竟意味着什么,目前学界并无定论。这篇论文换了一个切入角度:与其盯着神经网络这类"黑盒"追问为何难以解释,不如先搞清楚那些公认"好解释"的模型——线性模型、决策树、MARS 和广义加性模型(GAM)——到底好在哪里。作者借助哲学中关于"理解"的理论框架,逐一分析这四类模型让人"看得懂"的具体属性,例如参数可直接对应输入变量的权重、预测路径可视化、输出对输入的影响可分解追踪等。研究发现,不同模型带来可解释性的方式各有差异,并不存在一个放之四海而皆准的单一机制,但每种情形下"可解释到什么程度、为什么可解释"都可以被清晰地刻画出来。这项工作的意义在于:只有先把"可解释性"这个概念本身说清楚,才能真正评估和改进那些黑盒模型的透明化方法,为可解释 AI 的研究提供更扎实的概念基础。
原文 arXiv:2211.13617;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2211.13617v1