Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach
Weiyu Ma1,2, Qirui Mi1,2, Yongcheng Zeng1,2, Xue Yan1,2, Yuqiao Wu1,2, Runji Lin1,2, Haifeng Zhang,2,4124{}^{~{}~{}1,2,4}start_FLOATSUPERSCRIPT 1 , 2 , 4 end_FLOATSUPERSCRIPT, Jun Wang 33{}^{~{}3}start_FLOATSUPERSCRIPT 3 end_FLOATSUPERSCRIPT 1 Institute of Automation, Chinese Academy of Sciences, China 2 School of Artificial Intelligence, University of Chinese Academy of Sciences, China 3 Department of Computer Science, University College London, UK 4 Nanjing Artificial Intelligence Research of IA, China Corresponding to Haifeng and Jun Wang
Abstract
With the continued advancement of Large Language Models (LLMs) Agents in reasoning, planning, and decision-making, benchmarks have become crucial in evaluating these skills. However, there is a notable gap in benchmarks for real-time strategic decision-making. StarCraft II (SC2), with its complex and dynamic nature, serves as an ideal setting for such evaluations. To this end, we have developed TextStarCraft II, a specialized environment for assessing LLMs in real-time strategic scenarios within SC2. Addressing the limitations of traditional Chain of Thought (CoT) methods, we introduce the Chain of Summarization (CoS) method, enhancing LLMs’ capabilities in rapid and effective decision-making. Our key experiments included: 1. LLM Evaluation: Tested 10 LLMs in TextStarCraft II, most of them defeating LV5 build-in AI, showcasing effective strategy skills. 2. Commercial Model Knowledge: Evaluated four commercial models on SC2 knowledge; GPT-4 ranked highest by Grandmaster-level experts. 3. Human-AI Matches: Experimental results showed that fine-tuned LLMs performed on par with Gold-level players in real-time matches, demonstrating comparable strategic abilities.
中文速览
星际争霸II(StarCraft II)对实时战略决策和长期规划的要求极高,但目前缺乏专门用来评估大语言模型(LLM)在此类场景下能力的基准测试平台。为此,研究团队开发了TextStarCraft II——一个将星际争霸II的复杂游戏状态转化为文本格式的交互环境,让大语言模型能够通过自然语言指令来执行宏观战略决策。针对传统思维链(Chain of Thought, CoT)方法在处理长时序复杂信息时效率低下的问题,他们进一步提出了"摘要链"(Chain of Summarization, CoS)方法,通过单帧摘要与多帧摘要模块对游戏状态信息进行压缩与聚合,从而加快决策速度并提升全局理解能力。实验结果显示,测试的10个大语言模型中大多数能击败游戏内置五级AI,GPT-4在星际争霸知识问答评测中被大师级人类专家评为最强,而经过微调的开源模型甚至能与黄金段位(Gold-level)人类玩家相抗衡。这项工作首次为大语言模型在实时战略游戏中的推理与规划能力提供了系统性评测框架,对推动LLM走向更复杂的动态决策场景具有重要参考价值。
原文 arXiv:2312.11865;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2312.11865v3