Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Aravind Rajeswaran1∗ * These authors contributed equally to this work. Vikash Kumar1,2∗ Abhishek Gupta3 Giulia Vezzani4 John Schulman2 Emanuel Todorov1 Sergey Levine3
Abstract
Dexterous multi-fingered hands are extremely versatile and provide a generic way to perform a multitude of tasks in human-centric environments. However, effectively controlling them remains challenging due to their high dimensionality and large number of potential contacts. Deep reinforcement learning (DRL) provides a model-agnostic approach to control complex dynamical systems, but has not been shown to scale to high-dimensional dexterous manipulation. Furthermore, deployment of DRL on physical systems remains challenging due to sample inefficiency. Consequently, the success of DRL in robotics has thus far been limited to simpler manipulators and tasks. In this work, we show that model-free DRL can effectively scale up to complex manipulation tasks with a high-dimensional 24-DoF hand, and solve them from scratch in simulated experiments. Furthermore, with the use of a small number of human demonstrations, the sample complexity can be significantly reduced, which enables learning with sample sizes equivalent to a few hours of robot experience. The use of demonstrations result in policies that exhibit very natural movements and, surprisingly, are also substantially more robust. We d
中文速览
多指灵巧手(dexterous multi-fingered hand)虽然能完成人类日常所需的各种精细操作,但因其自由度高达24维、接触模式复杂,用深度强化学习(deep reinforcement learning, DRL)直接控制它一直被认为难以实现。研究团队在仿真环境中设计了物体搬运、手内操作、工具使用和开门四类代表性任务,证明无模型DRL确实能从零开始学会这些高难度操作;进一步地,他们通过虚拟现实收集少量人类示范,先用行为克隆做策略预训练,再结合策略梯度方法微调,显著压缩了所需样本量,使训练数据量等价于机器人在现实中工作几小时便可完成。加入示范后,学到的策略不仅更自然流畅,鲁棒性也大幅提升,表明人类示范中隐含的操作先验能引导强化学习走向更稳健的解。这项工作首次说明无模型DRL配合少量示范在高维灵巧手操作上切实可行,为未来在真实机器人上直接学习复杂操作技能打开了一扇门。
原文 arXiv:1709.10087;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1709.10087v2