arXiv:1712.01815 · 中英对照阅读
通过自我对弈与通用强化学习算法掌握国际象棋和将棋
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
中文速览
传统棋类程序依赖人类设计的评估函数和大量领域技巧,难以迁移到不同游戏,AlphaZero则试图只凭游戏规则和自我对弈学会下棋。它用深度神经网络同时预测下一步走法和局面胜负,再结合通用的蒙特卡洛树搜索(MCTS)反复自我对弈、更新模型,从随机初始化出发训练国际象棋、日本将棋和围棋。结果显示,AlphaZero在几小时内就超过了当时最强的Stockfish、Elmo和旧版AlphaGo Zero,且只搜索约千分之一的局面,仍能更有效地把计算集中在关键变化上。其重要性在于,它证明了不依赖人工棋理和专门程序设计的通用强化学习方法,也能在多个复杂棋类领域达到并超越顶尖水平。
摘要
国际象棋是人工智能发展史上研究最为广泛的领域。最强大的程序基于复杂搜索技术、领域特定适配以及手工设计的评估函数的组合,这些方法经过数十年人类专家的不断改进。相比之下,AlphaGo Zero 程序最近通过从自我对弈棋局中进行白板式学习的强化学习,在围棋领域达到了超人类水平。本文将这一方法推广为一种单一的 AlphaZero 算法,使其能够在多个具有挑战性的领域中以白板式学习的方式达到超人类水平。从随机对弈开始,除游戏规则外不具备任何领域知识,AlphaZero 在 24 小时内便在国际象棋、将棋(日本象棋)以及围棋中达到了超人类水平,并且在每个项目中都令人信服地击败了世界冠军级程序。
The game of chess is the most widely-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques, domain-specific adaptations, and handcrafted evaluation functions that have been refined by human experts over several decades. In contrast, the AlphaGo Zero program recently achieved superhuman performance in the game of Go, by tabula rasa reinforcement learning from games of self-play. In this paper, we generalise this approach into a single AlphaZero algorithm that can achieve, tabula rasa, superhuman performance in many challenging domains. Starting from random play, and given no domain knowledge except the game rules, AlphaZero achieved within 24 hours a superhuman level of play in the games of chess and shogi (Japanese chess) as well as Go, and convincingly defeated a world-champion program in each case.
术语表
- AlphaZero
- AlphaZero
- AlphaGo Zero
- AlphaGo Zero
- tabula rasa
- 白板式学习
- reinforcement learning
- 强化学习
- self-play
- 自我对弈
- superhuman performance
- 超人类水平
- domain-specific adaptation
- 领域特定适配
- handcrafted evaluation function
- 手工设计的评估函数
- alpha-beta search
- alpha-beta 搜索
- search tree
- 搜索树
- Monte-Carlo tree search (MCTS)
- 蒙特卡洛树搜索(MCTS)
- deep neural network
- 深度神经网络
- deep convolutional neural network
- 深度卷积神经网络
- policy network
- 策略网络
- value network
- 价值网络
- move probability
- 着法概率
- expected outcome
- 期望结果
- policy vector
- 策略向量
- visit count
- 访问次数
- gradient descent
- 梯度下降
- mean-squared error
- 均方误差
- cross-entropy loss
- 交叉熵损失
- L2 weight regularisation
- L2 权重正则化
- data augmentation
- 数据增强
- ensembling
- 集成
- Top Chess Engine Championship (TCEC)
- 顶级国际象棋引擎锦标赛(TCEC)
- Stockfish
- Stockfish
- Computer Shogi Association (CSA)
- 日本将棋协会(CSA)
- Elmo
- Elmo