Aha.
正在载入中英对照阅读…

arXiv:1712.01815 · 中英对照阅读

通过自我对弈与通用强化学习算法掌握国际象棋和将棋

Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

David Silver、Thomas Hubert、Julian Schrittwieser、Ioannis Antonoglou、Matthew Lai、Arthur Guez、Marc Lanctot、Laurent Sifre、Dharshan Kumaran、Thore Graepel、Timothy Lillicrap、Karen Simonyan、Demis Hassabis

中文速览

传统棋类程序依赖人类设计的评估函数和大量领域技巧,难以迁移到不同游戏,AlphaZero则试图只凭游戏规则和自我对弈学会下棋。它用深度神经网络同时预测下一步走法和局面胜负,再结合通用的蒙特卡洛树搜索(MCTS)反复自我对弈、更新模型,从随机初始化出发训练国际象棋、日本将棋和围棋。结果显示,AlphaZero在几小时内就超过了当时最强的Stockfish、Elmo和旧版AlphaGo Zero,且只搜索约千分之一的局面,仍能更有效地把计算集中在关键变化上。其重要性在于,它证明了不依赖人工棋理和专门程序设计的通用强化学习方法,也能在多个复杂棋类领域达到并超越顶尖水平。

摘要

国际象棋是人工智能发展史上研究最为广泛的领域。最强大的程序基于复杂搜索技术、领域特定适配以及手工设计的评估函数的组合,这些方法经过数十年人类专家的不断改进。相比之下,AlphaGo Zero 程序最近通过从自我对弈棋局中进行白板式学习的强化学习,在围棋领域达到了超人类水平。本文将这一方法推广为一种单一的 AlphaZero 算法,使其能够在多个具有挑战性的领域中以白板式学习的方式达到超人类水平。从随机对弈开始,除游戏规则外不具备任何领域知识,AlphaZero 在 24 小时内便在国际象棋、将棋(日本象棋)以及围棋中达到了超人类水平,并且在每个项目中都令人信服地击败了世界冠军级程序。

The game of chess is the most widely-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques, domain-specific adaptations, and handcrafted evaluation functions that have been refined by human experts over several decades. In contrast, the AlphaGo Zero program recently achieved superhuman performance in the game of Go, by tabula rasa reinforcement learning from games of self-play. In this paper, we generalise this approach into a single AlphaZero algorithm that can achieve, tabula rasa, superhuman performance in many challenging domains. Starting from random play, and given no domain knowledge except the game rules, AlphaZero achieved within 24 hours a superhuman level of play in the games of chess and shogi (Japanese chess) as well as Go, and convincingly defeated a world-champion program in each case.

术语表

AlphaZero
AlphaZero
AlphaGo Zero
AlphaGo Zero
tabula rasa
白板式学习
reinforcement learning
强化学习
self-play
自我对弈
superhuman performance
超人类水平
domain-specific adaptation
领域特定适配
handcrafted evaluation function
手工设计的评估函数
alpha-beta search
alpha-beta 搜索
search tree
搜索树
Monte-Carlo tree search (MCTS)
蒙特卡洛树搜索(MCTS)
deep neural network
深度神经网络
deep convolutional neural network
深度卷积神经网络
policy network
策略网络
value network
价值网络
move probability
着法概率
expected outcome
期望结果
policy vector
策略向量
visit count
访问次数
gradient descent
梯度下降
mean-squared error
均方误差
cross-entropy loss
交叉熵损失
L2 weight regularisation
L2 权重正则化
data augmentation
数据增强
ensembling
集成
Top Chess Engine Championship (TCEC)
顶级国际象棋引擎锦标赛(TCEC)
Stockfish
Stockfish
Computer Shogi Association (CSA)
日本将棋协会(CSA)
Elmo
Elmo