Causal Parrots: Large Language Models May Talk Causality But Are Not Causal
Matej Zečević* Computer Science Department, TU Darmstadt, Germany Moritz Willig* Computer Science Department, TU Darmstadt, Germany Devendra Singh Dhami Computer Science Department, TU Darmstadt, Germany Hessian Center for AI (hessian.AI), Germany Kristian Kersting Computer Science Department, TU Darmstadt, Germany Centre for Cognitive Science, TU Darmstadt, Germany Hessian Center for AI (hessian.AI), Germany German Research Center for Artificial Intelligence (DFKI), Germany
Abstract
Some argue scale is all what is needed to achieve AI, covering even causal models. We make it clear that large language models (LLMs) cannot be causal and give reason onto why sometimes we might feel otherwise. To this end, we define and exemplify a new subgroup of Structural Causal Model (SCM) that we call meta SCM which encode causal facts about other SCM within their variables. We conjecture that in the cases where LLM succeed in doing causal inference, underlying was a respective meta SCM that exposed correlations between causal facts in natural language on whose data the LLM was ultimately trained. If our hypothesis holds true, then this would imply that LLMs are like parrots in that they simply recite the causal knowledge embedded in the data. Our empirical analysis provides favoring evidence that current LLMs are even weak ‘causal parrots.’
中文速览
大型语言模型(LLM)究竟懂不懂因果推理,还是只会"鹦鹉学舌"?作者引入"元结构因果模型"(meta Structural Causal Model,meta SCM)这一新概念来回答这个问题——meta SCM 把关于因果关系的事实本身编码为变量,使得因果知识以相关性的形式暗含在自然语言文本中。据此,作者提出核心假说:LLM 之所以有时能答对因果问题,并非真正理解因果机制,而是因为训练数据(如维基百科)中已经包含了人类总结好的因果事实,模型只是学会了重复这些"因果相关性",本质上是"因果鹦鹉"(causal parrots)。实验结果进一步表明,现有 LLM 连"因果鹦鹉"都做得不够好,在超出训练数据覆盖范围的因果问题上表现明显下滑。这项工作的意义在于:它从理论上说明了单纯扩大模型规模和数据量并不能让 LLM 真正具备因果推理能力,为"规模即一切"的论断提供了有力的反驳证据。
原文 arXiv:2308.13067;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2308.13067v1