Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration
Xavier Puig1 Tianmin Shu1 Shuang Li1 Zilin Wang2 Yuan-Hong Liao3,5 Joshua B. Tenenbaum1 Sanja Fidler3,4,5 Antonio Torralba1 1Massachusetts Institute of Technology 2ETH Zurich 3University of Toronto 4NVIDIA 5Vector Institute
Abstract
In this paper, we introduce Watch-And-Help (WAH), a challenge for testing social intelligence in agents. In WAH, an AI agent needs to help a human-like agent perform a complex household task efficiently. To succeed, the AI agent needs to i) understand the underlying goal of the task by watching a single demonstration of the human-like agent performing the same task (social perception), and ii) coordinate with the human-like agent to solve the task in an unseen environment as fast as possible (human-AI collaboration). For this challenge, we build VirtualHome-Social, a multi-agent household environment, and provide a benchmark including both planning and learning based baselines. We evaluate the performance of AI agents with the human-like agent as well as with real humans using objective metrics and subjective user ratings. Experimental results demonstrate that the proposed challenge and virtual environment enable a systematic evaluation on the important aspects of machine social intelligence at scale.111Code and documentation for the VirtualHome-Social environment are available at https://virtual-home.org. Code and data for the WAH challenge are available at https://github.com/xavi
中文速览
让AI真正学会"帮人干活"是个难题——它不仅要看懂别人在做什么,还要主动配合对方把事情做完。Watch-And-Help(WAH)挑战正是为此设计的:AI智能体(Bob)先观看另一个类人智能体(Alice)完成一项家务任务的单次示范,推断出她的目标,然后在一个全新的环境里与Alice协作,用尽可能少的步骤帮她完成同样的任务。为了支撑这个挑战,作者构建了多智能体仿真平台VirtualHome-Social,其中内置了基于蒙特卡洛树搜索的类人规划智能体,并支持真实人类参与交互。实验结果表明,现有的规划和深度强化学习基线在目标推断与协作泛化两个方面都还存在明显短板,证明WAH能有效衡量机器社会智能的核心能力。这项工作的价值在于,它首次提供了一个可大规模、系统化评估AI社会感知与人机协作能力的真实场景测试床,为未来开发真正能理解并帮助人类的AI奠定了基础。
原文 arXiv:2010.09890;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2010.09890v2