Towards Generalization and Simplicity in Continuous Control
Aravind Rajeswaran Kendall Lowrey Emanuel Todorov Sham Kakade[0.2cm] University of Washington Seattle[0.2cm] {\{ aravraj, klowrey, todorov, sham }\} @ cs.washington.edu
Abstract
This work shows that policies with simple linear and RBF parameterizations can be trained to solve a variety of widely studied continuous control tasks, including the OpenAI gym benchmarks. The performance of these trained policies are competitive with state of the art results, obtained with more elaborate parameterizations such as fully connected neural networks. Furthermore, the standard training and testing scenarios for these tasks are shown to be very limited and prone to over-fitting, thus giving rise to only trajectory-centric policies. Training with a diverse initial state distribution induces more global policies with better generalization. This allows for interactive control scenarios where the system recovers from large on-line perturbations; as shown in the supplementary video.
原文 arXiv:1703.02660;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1703.02660v2