BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning
Maxime Chevalier-Boisvert Mila, Université de Montréal、Dzmitry Bahdanau Mila, Université de Montréal AdeptMind Scholar Element AI \ANDSalem Lahlou Mila, Université de Montréal、Lucas Willems École Normale Supérieure, Paris、Chitwan Saharia IIT Bombay \ANDThien Huu Nguyen University of Oregon、Yoshua Bengio Mila, Université de Montréal CIFAR Senior Fellow Equal contribution.Work done during an internship at Mila.Work done during a post-doc at Mila.
Abstract
Allowing humans to interactively train artificial agents to understand language instructions is desirable for both practical and scientific reasons. Though, given the lack of sample efficiency in current learning methods, reaching this goal may require substantial research efforts. We introduce the BabyAI research platform, with the goal of supporting investigations towards including humans in the loop for grounded language learning. The BabyAI platform comprises an extensible suite of 19 levels of increasing difficulty. Each level gradually leads the agent towards acquiring a combinatorially rich synthetic language, which is a proper subset of English. The platform also provides a hand-crafted bot agent, which simulates a human teacher. We report estimated amount of supervision required for training neural reinforcement and behavioral-cloning agents on some BabyAI levels. We put forward strong evidence that current deep learning methods are not yet sufficiently sample-efficient in the context of learning a language with compositional properties.
中文速览
让人类通过语言指令教会人工智能在真实环境中行动,最大的难题是如何用尽量少的人类示范和反馈学会组合式语言。研究者搭建了 BabyAI 平台:在可部分观察、能开门和搬动物体的二维网格世界中,设计了由易到难的 19 个任务和一套可组合的合成语言,并用能随时示范和指导的机器人模拟人类教师,再测试行为克隆、强化学习、预训练和交互式模仿学习等方法。实验显示,现有深度学习方法要把这些任务学好仍需大量示范或交互,面对具有组合结构的语言时样本效率明显不够,预训练和互动教学虽有帮助但尚未根本解决问题。这个平台因此既提供了统一的研究环境和基准,也明确指出了让真实人类高效教会智能体理解语言前最需要突破的瓶颈。
原文 arXiv:1810.08272;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1810.08272v4