Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
David Silver Thomas Hubert Julian Schrittwieser Ioannis Antonoglou Matthew Lai Affiliation: DeepMind, 6 Pancras Square, London N1C 4AG.∗These authors contributed equally to this work. Arthur Guez Affiliation: DeepMind, 6 Pancras Square, London N1C 4AG.∗These authors contributed equally to this work. Marc Lanctot Affiliation: DeepMind, 6 Pancras Square, London N1C 4AG.∗These authors contributed equally to this work. Laurent Sifre Dharshan Kumaran Affiliation: DeepMind, 6 Pancras Square, London N1C 4AG.∗These authors contributed equally to this work. Thore Graepel Affiliation: DeepMind, 6 Pancras Square, London N1C 4AG.∗These authors contributed equally to this work. Timothy Lillicrap Karen Simonyan Affiliation: DeepMind, 6 Pancras Square, London N1C 4AG.∗These authors contributed equally to this work. Demis Hassabis Affiliation: DeepMind, 6 Pancras Square, London N1C 4AG.∗These authors contributed equally to this work.
Abstract
The game of chess is the most widely-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques, domain-specific adaptations, and handcrafted evaluation functions that have been refined by human experts over several decades. In contrast, the AlphaGo Zero program recently achieved superhuman performance in the game of Go, by tabula rasa reinforcement learning from games of self-play. In this paper, we generalise this approach into a single AlphaZero algorithm that can achieve, tabula rasa, superhuman performance in many challenging domains. Starting from random play, and given no domain knowledge except the game rules, AlphaZero achieved within 24 hours a superhuman level of play in the games of chess and shogi (Japanese chess) as well as Go, and convincingly defeated a world-champion program in each case.
中文速览
传统棋类程序依赖人类设计的评估函数和大量领域技巧,难以迁移到不同游戏,AlphaZero则试图只凭游戏规则和自我对弈学会下棋。它用深度神经网络同时预测下一步走法和局面胜负,再结合通用的蒙特卡洛树搜索(MCTS)反复自我对弈、更新模型,从随机初始化出发训练国际象棋、日本将棋和围棋。结果显示,AlphaZero在几小时内就超过了当时最强的Stockfish、Elmo和旧版AlphaGo Zero,且只搜索约千分之一的局面,仍能更有效地把计算集中在关键变化上。其重要性在于,它证明了不依赖人工棋理和专门程序设计的通用强化学习方法,也能在多个复杂棋类领域达到并超越顶尖水平。
原文 arXiv:1712.01815;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1712.01815v1