On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
Dan Ley、Sree Harsha Tanneru、Chirag Agarwal、Himabindu Lakkaraju Harvard University Cambridge, MA 02138 Equal Contribution. Correspondence to Sree Harsha Tanneru
Abstract
As Large Language Models (LLMs) are increasingly being employed in real-world applications in critical domains such as healthcare, it is important to ensure that the Chain-of-Thought (CoT) reasoning generated by these models faithfully captures their underlying behavior. While LLMs are known to generate CoT reasoning that is appealing to humans, prior studies have shown that these explanations do not accurately reflect the actual behavior of the underlying LLMs. In this work, we explore the promise of three broad approaches commonly employed to steer the behavior of LLMs to enhance the faithfulness of the CoT reasoning generated by LLMs: in-context learning, fine-tuning, and activation editing. Specifically, we introduce novel strategies for in-context learning, fine-tuning, and activation editing aimed at improving the faithfulness of the CoT reasoning. We then carry out extensive empirical analyses with multiple benchmark datasets to explore the promise of these strategies. Our analyses indicate that these strategies offer limited success in improving the faithfulness of the CoT reasoning, with only slight performance enhancements in controlled scenarios. Activation editing demon
中文速览
大型语言模型(LLM)在医疗等高风险场景中越来越多地被使用,但它们生成的思维链(Chain-of-Thought, CoT)推理过程往往只是"看起来合理",并不真实反映模型内部的实际决策逻辑——这一忠实性(faithfulness)问题严重制约了人们对模型的信任与监督。为此,研究者系统测试了三类常见的模型行为干预手段:上下文学习(in-context learning,在推理时提供忠实推理示例)、微调(fine-tuning,用忠实样本更新模型参数)以及激活编辑(activation editing,定位并沿"忠实方向"修改注意力头的内部激活),并针对每类方法设计了多种具体策略。在多个推理和问答基准数据集上的大量实验表明,三类方法的效果均十分有限:激活编辑几乎没有带来改善,微调和上下文学习仅在受控场景下有边际提升,且无法泛化到不同任务。这一结果揭示了让LLM生成真正忠实的推理链是一个内在困难的问题,现有技术手段远远不够,呼唤全新的方法论来深入理解模型内部机制。
原文 arXiv:2406.10625;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2406.10625v2