A Game-Theoretic and Introspective Framework for Interpreting Neural Predictions
First Author Affiliation: Affiliation / Address line 1 Affiliation: Affiliation / Address line 2 Affiliation: Affiliation / Address line 3 Email: Second Author Affiliation: Affiliation / Address line 1 Affiliation: Affiliation / Address line 2 Affiliation: Affiliation / Address line 3 Email:
Abstract
Many recent self-explainable models for NLP adopt a two-step extractive framework where the first model generates a subset of the input as the explanation and the second model makes prediction over this explanation. In this paper, we first point out that all these works fall to a collaborative game formulation thus may suffer from unexpected behaviors, including (1) the two models can simply communicate with their own specific code on class labels, making the extracted explanations less meaningful; and (2) training of the first-step model may be difficult to capture sufficient information especially with less discriminative supervision. To deal with these problems, we propose a new three-player game framework to filter out the undesirable communications, and a new introspective model to encourage the explanation generation to be more aware of potential predictions. The new framework theoretically guarantees the good properties of the generated explanations, as well as gives improved empirical results on text classification tasks.
原文 arXiv:1910.13294;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1910.13294v2