Latent Offline Model-Based Policy Optimization
Albert Author1 and Bernard D. Researcher2 *This work was not supported by any organization1Albert Author is with Faculty of Electrical Engineering, Mathematics and Computer Science, University of Twente, 7500 AE Enschede, The Netherlands D. Researcheris with the Department of Electrical Engineering, Wright State University, Dayton, OH 45435, USA
Abstract
Offline reinforcement learning (RL) has shown promise in learning policies from a prerecorded dataset without any interaction with the environment, and has the potential to significantly improve the scalability and safety of policy learning. Recent advances in offline model-based RL have further improved previous offline model-free approaches by enabling greater generalization to new states. These advantages make offline model-based RL particularly appealing for learning complex skills from raw sensor observations, such as images, since learning visuomotor policies using deep RL methods is much more sample-inefficient and unsafe (e.g. failure to generalize to lighting changes) than learning from low-dimensional state inputs. Hence, learning vision-based tasks end-to-end from offline data becomes crucial for robotics and control. However, offline model-based RL from visual inputs is challenging because the offline model-based RL problem critically relies on accurate uncertainty quantification of the model’s predictions to avoid falling off the data distribution and estimating uncertainty of high-dimensional visual dynamics models such as learning an ensemble of video prediction mode
原文 arXiv:2012.11547;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2012.11547v1