Acme: A Research Framework for Distributed Reinforcement Learning
Matthew W. Hoffman*†, Bobak Shahriari*†, John Aslanides†, Gabriel Barth-Maron† DeepMind Nikola Momchev*, Danila Sinopalnikov*, Piotr Stańczyk*, Sabela Ramos, Anton Raichuk, Damien Vincent Google Research, Brain Team Léonard Hussenot*, Robert Dadashi*, Gabriel Dulac-Arnold, Manu Orsini, Alexis Jacq, Johan Ferret, Nino Vieillard, Seyed Kamyar Seyed Ghasemipour, Sertan Girgin, Olivier Pietquin Google Research, Brain Team Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, Sarah Henderson, Abe Friesen, Ruba Haroun‡, Alex Novikov, Sergio Gómez Colmenarejo, Serkan Cabi, Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Andrew Cowie, Ziyu Wang‡, Bilal Piot, Nando de Freitas DeepMind, ‡Work done while at DeepMind *Core Contributor †Original Author
Abstract
Deep reinforcement learning (RL) has led to many recent and groundbreaking advances. However, these advances have often come at the cost of both increased scale in the underlying architectures being trained as well as increased complexity of the RL algorithms used to train them. These increases have in turn made it more difficult for researchers to rapidly prototype new ideas or reproduce published RL algorithms. To address these concerns this work describes Acme, a framework for constructing novel RL algorithms that is specifically designed to enable agents that are built using simple, modular components that can be used at various scales of execution. While the primary goal of Acme is to provide a framework for algorithm development, a secondary goal is to provide simple reference implementations of important or state-of-the-art algorithms. These implementations serve both as a validation of our design decisions as well as an important contribution to reproducibility in RL research. In this work we describe the major design decisions made within Acme and give further details as to how its components can be used to implement various algorithms. Our experiments provide baselines fo
中文速览
深度强化学习(RL)算法越来越强大,但代价是模型规模越来越大、算法越来越复杂,导致研究者很难快速验证新想法或复现已发表的成果。为此,DeepMind 推出了 Acme 框架,把 RL 智能体拆解为若干简洁、可复用的模块——从网络结构、损失函数、策略,到演员(actor)、学习器(learner)、经验回放缓冲区,再到完整的训练循环与日志系统——让研究者既能在单机同步模式下轻松调试,也能无缝扩展到大规模分布式训练,而无需重写整套代码。框架内置了包括在线 RL、离线 RL(offline RL)及模仿学习(imitation learning)在内的多类前沿算法的参考实现,实验结果验证了这些实现在多个常见基准上的竞争力,同时展示了同一套代码如何平滑地扩展到更复杂的大规模环境。Acme 的意义在于,它让"写出可读、可复现、又能跑出顶尖性能"的 RL 代码不再互相矛盾,从而降低整个领域的研究门槛并推动结果可复现性。
原文 arXiv:2006.00979;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2006.00979v2