VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Wenlong Huang Affiliation: Stanford University Chen Wang Affiliation: Stanford University Ruohan Zhang Affiliation: Stanford University Yunzhu Li Affiliation: Stanford University Affiliation: University of Illinois Urbana-Champaign Jiajun Wu Affiliation: Stanford University Li Fei-Fei Affiliation: Stanford University
Abstract
Large language models (LLMs) are shown to possess a wealth of actionable knowledge that can be extracted for robot manipulation in the form of reasoning and planning. Despite the progress, most still rely on pre-defined motion primitives to carry out the physical interactions with the environment, which remains a major bottleneck. In this work, we aim to synthesize robot trajectories, i.e., a dense sequence of 6-DoF end-effector waypoints, for a large variety of manipulation tasks given an open-set of instructions and an open-set of objects. We achieve this by first observing that LLMs excel at inferring affordances and constraints given a free-form language instruction. More importantly, by leveraging their code-writing capabilities, they can interact with a vision-language model (VLM) to compose 3D value maps to ground the knowledge into the observation space of the agent. The composed value maps are then used in a model-based planning framework to zero-shot synthesize closed-loop robot trajectories with robustness to dynamic perturbations. We further demonstrate how the proposed framework can benefit from online experiences by efficiently learning a dynamics model for scenes tha
原文 arXiv:2307.05973;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2307.05973v2