Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
Chirag Agarwal Sree Harsha Tanneru Himabindu Lakkaraju
Abstract
Large Language Models (LLMs) are deployed as powerful tools for several natural language processing (NLP) applications. Recent works show that modern LLMs can generate self-explanations (SEs), which elicit their intermediate reasoning steps for explaining their behavior. Self-explanations have seen widespread adoption owing to their conversational and plausible nature. However, there is little to no understanding of their faithfulness. In this work, we discuss the dichotomy between faithfulness and plausibility in SEs generated by LLMs. We argue that while LLMs are adept at generating plausible explanations – seemingly logical and coherent to human users – these explanations do not necessarily align with the reasoning processes of the LLMs, raising concerns about their faithfulness. We highlight that the current trend towards increasing the plausibility of explanations, primarily driven by the demand for user-friendly interfaces, may come at the cost of diminishing their faithfulness. We assert that the faithfulness of explanations is critical in LLMs employed for high-stakes decision-making. Moreover, we emphasize the need for a systematic characterization of faithfulness-plausibi
中文速览
大语言模型(LLM)在生成"自我解释"(self-explanation,SE)时,往往能给出听起来合理、逻辑连贯的推理过程,但这些解释未必真实反映模型内部的决策机制——这正是"似真性"(plausibility)与"忠实性"(faithfulness)之间的根本矛盾。作者系统梳理了链式思维(chain-of-thought)、词元重要性、反事实解释等主流自我解释方法,指出当前业界为追求用户友好的交互体验而不断提升解释的似真性,却在无意中牺牲了其对模型真实推理过程的准确描述。通过对医疗、法律、金融等高风险场景的分析,论文表明一旦解释只是"听起来有道理"而非"真实可信",就可能误导决策者并带来严重后果。研究呼吁学界针对不同应用场景明确似真性与忠实性的需求边界,并开发新方法切实提升自我解释的忠实性,从而让LLM在高风险领域得到更透明、更可靠的部署。
原文 arXiv:2402.04614;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2402.04614v3