Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Tomer D. Ullman Department of Psychology Harvard University Cambridge, MA, 02138
Abstract
Intuitive psychology is a pillar of common-sense reasoning. The replication of this reasoning in machine intelligence is an important stepping-stone on the way to human-like artificial intelligence. Several recent tasks and benchmarks for examining this reasoning in Large-Large Models have focused in particular on belief attribution in Theory-of-Mind tasks. These tasks have shown both successes and failures. We consider in particular a recent purported success case (kosinski2023theory, ), and show that small variations that maintain the principles of ToM turn the results on their head. We argue that in general, the zero-hypothesis for model evaluation in intuitive psychology should be skeptical, and that outlying failure cases should outweigh average success rates. We also consider what possible future successes on Theory-of-Mind tasks by more powerful LLMs would mean for ToM tasks with people.
中文速览
大型语言模型(LLM)是否真正具备"心理理论"(Theory of Mind,ToM)——即理解他人信念与心理状态的能力——是当前AI研究的热点争议。针对Kosinski等人声称GPT-3.5已达9岁儿童ToM水平的论文,本文通过对经典ToM测试场景做微小但合理的改动来检验该结论的稳健性:例如让袋子变成透明的、让人物不识字因而无法读懂标签、或让人物事先被可信朋友告知真相——这些改动在人类看来完全不影响正确判断,但GPT-3.5仍然给出错误答案。结果表明,模型的"成功"高度依赖表面文本模式,而非真正理解他人的信念形成过程,少数关键失败案例足以否定"模型掌握了ToM"的结论,正如一道乘法算错就能说明机器没有真正学会乘法算法。这项研究提醒我们:评估AI的认知能力应以怀疑论为默认立场,不能用平均正确率掩盖根本性的理解缺失,否则既误判了机器能力,也可能动摇人类ToM测量工具本身的公信力。
原文 arXiv:2302.08399;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2302.08399v5