Large Language Models Cannot Self-Correct Reasoning Yet
Jie Huang1,212{}^{1,2}start_FLOATSUPERSCRIPT 1 , 2 end_FLOATSUPERSCRIPT Xinyun Chen11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT††Swaroop Mishra11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Huaixiu Steven Zheng11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Adams Wei Yu11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Xinying Song11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Denny Zhou11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT 11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPTGoogle DeepMind 22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPTUniversity of Illinois at Urbana-Champaign {xinyunchen, Equal contribution.
Abstract
Large Language Models (LLMs) have emerged as a groundbreaking technology with their unparalleled text generation capabilities across various applications. Nevertheless, concerns persist regarding the accuracy and appropriateness of their generated content. A contemporary methodology, self-correction, has been proposed as a remedy to these issues. Building upon this premise, this paper critically examines the role and efficacy of self-correction within LLMs, shedding light on its true potential and limitations. Central to our investigation is the notion of intrinsic self-correction, whereby an LLM attempts to correct its initial responses based solely on its inherent capabilities, without the crutch of external feedback. In the context of reasoning, our research indicates that LLMs struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction. Drawing from these insights, we offer suggestions for future research and practical applications in this field.
中文速览
大型语言模型(LLM)能否在没有外部反馈的情况下,仅凭自身能力发现并纠正推理错误?研究者针对这一"内在自我纠错"(intrinsic self-correction)能力展开了系统检验,在数学推理、常识问答和多跳问答等多个基准上测试了GPT-3.5、GPT-4、Llama-2等主流模型。实验结果令人警醒:在没有外部标准答案(oracle label)指引的情况下,自我纠错不仅没有提升模型准确率,反而在所有测试场景中都让性能下滑,根本原因在于模型无法可靠判断自己的回答是否正确,反而更容易把本来正确的答案改错。进一步分析还发现,此前声称自我纠错有效的研究,要么暗中借用了标准答案来引导纠错循环,要么用了较差的初始提示词使得后续纠错显得有效,一旦控制这些变量,优势便消失殆尽;多智能体辩论(multi-agent debate)同样不比自洽性采样(self-consistency)更优。这一发现对当前社区中对LLM自我纠错能力的乐观预期提出了重要质疑,提示研究者需要更严格的评估框架和真正有效的纠错机制。
原文 arXiv:2310.01798;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2310.01798v2