QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
Dmitry Kalashnikov1, Alex Irpan1, Peter Pastor2, Julian Ibarz1, Alexander Herzog2, Eric Jang1, Deirdre Quillen3, Ethan Holly1, Mrinal Kalakrishnan2, Vincent Vanhoucke1, Sergey Levine1,3 {dkalashnikov, alexirpan, julianibarz, ejang, eholly, vanhoucke, {peterpastor, alexherzog,
Abstract
In this paper, we study the problem of learning vision-based dynamic manipulation skills using a scalable reinforcement learning approach. We study this problem in the context of grasping, a longstanding challenge in robotic manipulation. In contrast to static learning behaviors that choose a grasp point and then execute the desired grasp, our method enables closed-loop vision-based control, whereby the robot continuously updates its grasp strategy based on the most recent observations to optimize long-horizon grasp success. To that end, we introduce QT-Opt, a scalable self-supervised vision-based reinforcement learning framework that can leverage over 580k real-world grasp attempts to train a deep neural network Q-function with over 1.2M parameters to perform closed-loop, real-world grasping that generalizes to 96% grasp success on unseen objects. Aside from attaining a very high success rate, our method exhibits behaviors that are quite distinct from more standard grasping systems: using only RGB vision-based perception from an over-the-shoulder camera, our method automatically learns regrasping strategies, probes objects to find the most effective grasps, learns to reposition ob
中文速览
用深度强化学习让机器人通过视觉实时"看着抓",而不是"看一眼、闭眼抓"——这是这项工作的核心出发点。研究团队提出了一个名为 QT-Opt 的可扩展自监督强化学习框架,利用7台真实机器人收集的58万次抓取尝试来训练一个参数量达120万的深度神经网络Q函数,整个过程仅凭肩部单目RGB摄像头的图像作为输入,无需深度传感器或人工标注。实验结果表明,该系统在从未见过的新物体上达到了96%的抓取成功率,并自发涌现出重新抓取、试探性拨动、非抓握式预操作等复杂行为,能在遭受外部干扰时动态纠偏。这项工作的意义在于,它证明了通过大规模离线数据与少量在线微调相结合的方式,深度强化学习完全可以在真实世界中习得泛化能力强、策略丰富的闭环视觉操控技能,为通用机器人操作迈出了重要一步。
原文 arXiv:1806.10293;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1806.10293v3