Open-Ended Learning Leads to Generally Capable Agents
Open-Ended Learning Team* Affiliation: DeepMind, London, UK Adam Stooke Affiliation: DeepMind, London, UK Anuj Mahajan Affiliation: DeepMind, London, UK Catarina Barros Affiliation: DeepMind, London, UK Charlie Deck Affiliation: DeepMind, London, UK Jakob Bauer Affiliation: DeepMind, London, UK Jakub Sygnowski Affiliation: DeepMind, London, UK Maja Trebacz Affiliation: DeepMind, London, UK Max Jaderberg Affiliation: DeepMind, London, UK Michael Mathieu Affiliation: DeepMind, London, UK Nat McAleese Affiliation: DeepMind, London, UK Nathalie Bradley-Schmieg Affiliation: DeepMind, London, UK Nathaniel Wong Affiliation: DeepMind, London, UK Nicolas Porcel Affiliation: DeepMind, London, UK Roberta Raileanu Affiliation: DeepMind, London, UK Steph Hughes-Fitt Affiliation: DeepMind, London, UK Valentin Dalibard Affiliation: DeepMind, London, UK Wojciech Marian Czarnecki Affiliation: DeepMind, London, UK
Abstract
Artificial agents have achieved great success in individual challenging simulated environments, mastering the particular tasks they were trained for, with their behaviour even generalising to maps and opponents that were never encountered in training. In this work we create agents that can perform well beyond a single, individual task, that exhibit much wider generalisation of behaviour to a massive, rich space of challenges. We define a universe of tasks within an environment domain and demonstrate the ability to train agents that are generally capable across this vast space and beyond. The environment is natively multi-agent, spanning the continuum of competitive, cooperative, and independent games, which are situated within procedurally generated physical 3D worlds. The resulting space is exceptionally diverse in terms of the challenges posed to agents, and as such, even measuring the learning progress of an agent is an open research problem. We propose an iterative notion of improvement between successive generations of agents, rather than seeking to maximise a singular objective, allowing us to quantify progress despite tasks being incomparable in terms of achievable rewards.
原文 arXiv:2107.12808;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2107.12808v2