Can Large Language Models Serve as Rational Players in Game Theory? A Systematic Analysis
Caoyun Fan, Jindou Chen, Yaohui Jin, Hao He These authors contributed equally. Corresponding author.
Abstract
Game theory, as an analytical tool, is frequently utilized to analyze human behavior in social science research. With the high alignment between the behavior of Large Language Models (LLMs) and humans, a promising research direction is to employ LLMs as substitutes for humans in game experiments, enabling social science research. However, despite numerous empirical researches on the combination of LLMs and game theory, the capability boundaries of LLMs in game theory remain unclear. In this research, we endeavor to systematically analyze LLMs in the context of game theory. Specifically, rationality, as the fundamental principle of game theory, serves as the metric for evaluating players’ behavior — building a clear desire, refining belief about uncertainty, and taking optimal actions. Accordingly, we select three classical games (dictator game, Rock-Paper-Scissors, and ring-network game) to analyze to what extent LLMs can achieve rationality in these three aspects. The experimental results indicate that even the current state-of-the-art LLM (GPT-4) exhibits substantial disparities compared to humans in game theory. For instance, LLMs struggle to build desires based on uncommon pref
中文速览
大型语言模型(LLM)能否真正像人类一样参与博弈论实验,目前学界缺乏系统性评估。研究者以博弈论中"理性玩家"的三个核心特征——建立清晰欲望、形成对不确定性的信念、据此采取最优行动——为框架,分别用独裁者博弈、石头剪刀布和环形网络博弈对GPT-3、GPT-3.5和GPT-4进行测试。结果发现:面对常见偏好时LLM尚能做出一致选择,但遇到非主流偏好(如利他主义)时表现大幅下滑;在石头剪刀布中,模型普遍无法从简单的历史规律中提炼出有效信念;而在需要综合欲望与信念做决策的环形网络博弈中,即便给予明确提示,LLM仍会忽略或擅自修改已推理出的信念。这些发现表明,就连当前最先进的GPT-4与人类理性玩家之间仍存在显著差距,将LLM直接用作社会科学博弈实验中的"人类替代品"需要更为审慎的态度。
原文 arXiv:2312.05488;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2312.05488v2