Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models
Ted Xiao1,*1{}^{1,*}start_FLOATSUPERSCRIPT 1 , * end_FLOATSUPERSCRIPT Harris Chan1,2,*12{}^{1,2,*}start_FLOATSUPERSCRIPT 1 , 2 , * end_FLOATSUPERSCRIPT Pierre Sermanet11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Ayzaan Wahid11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Anthony Brohan11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Karol Hausman11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Sergey Levine11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT Jonathan Tompson11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT *{}^{*}start_FLOATSUPERSCRIPT * end_FLOATSUPERSCRIPT Equal contribution 11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPTRobotics at Google 22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPTUniversity of Toronto Project website: https://instructionaugmentation.github.io
Abstract
Robotic manipulation policies that follow natural language instructions are typically trained from corpora of robot-language data that were either collected with specific tasks in mind or expensively relabeled by humans with varied language descriptions in hindsight. Recently, large-scale pretrained vision-language models (VLMs) have been applied to robotics for learning representations and scene descriptors. Can these pretrained models serve as automatic labelers for robot data, effectively importing Internet-scale knowledge into existing datasets with limited ground truth annotations? For example, if the original annotations contained templated task descriptions such as “pick apple”, a pretrained VLM-based labeler could significantly expand the number of semantic concepts available in the data and introduce spatial concepts such as “the apple on the right side of the table” or alternative phrasings such as “the red colored fruit”. To accomplish this, we introduce Data-driven Instruction Augmentation for Language-conditioned control (DIAL): we utilize semi-supervised language labels to propagate CLIP’s semantic knowledge onto large datasets of unlabeled demonstration data, from wh
原文 arXiv:2211.11736;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2211.11736v3