Muesli: Combining Improvements in Policy Optimization
Matteo Hessel Affiliation: DeepMind, London, UK Correspondence to: Ivo Danihelka Affiliation: DeepMind, London, UK Affiliation: University College London Correspondence to: Fabio Viola Affiliation: DeepMind, London, UK Arthur Guez Affiliation: DeepMind, London, UK Simon Schmitt Affiliation: DeepMind, London, UK Laurent Sifre Affiliation: DeepMind, London, UK Theophane Weber Affiliation: DeepMind, London, UK David Silver Affiliation: DeepMind, London, UK Affiliation: University College London Hado van Hasselt Affiliation: DeepMind, London, UK Correspondence to:
Abstract
We propose a novel policy update that combines regularized policy optimization with model learning as an auxiliary loss. The update (henceforth Muesli) matches MuZero’s state-of-the-art performance on Atari. Notably, Muesli does so without using deep search: it acts directly with a policy network and has computation speed comparable to model-free baselines. The Atari results are complemented by extensive ablations, and by additional results on continuous control and 9x9 Go.
原文 arXiv:2104.06159;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2104.06159v2