BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning
Maxime Chevalier-Boisvert Mila, Université de Montréal、Dzmitry Bahdanau Mila, Université de Montréal AdeptMind Scholar Element AI \ANDSalem Lahlou Mila, Université de Montréal、Lucas Willems École Normale Supérieure, Paris、Chitwan Saharia IIT Bombay \ANDThien Huu Nguyen University of Oregon、Yoshua Bengio Mila, Université de Montréal CIFAR Senior Fellow Equal contribution.Work done during an internship at Mila.Work done during a post-doc at Mila.
Abstract
Allowing humans to interactively train artificial agents to understand language instructions is desirable for both practical and scientific reasons. Though, given the lack of sample efficiency in current learning methods, reaching this goal may require substantial research efforts. We introduce the BabyAI research platform, with the goal of supporting investigations towards including humans in the loop for grounded language learning. The BabyAI platform comprises an extensible suite of 19 levels of increasing difficulty. Each level gradually leads the agent towards acquiring a combinatorially rich synthetic language, which is a proper subset of English. The platform also provides a hand-crafted bot agent, which simulates a human teacher. We report estimated amount of supervision required for training neural reinforcement and behavioral-cloning agents on some BabyAI levels. We put forward strong evidence that current deep learning methods are not yet sufficiently sample-efficient in the context of learning a language with compositional properties.
中文速览
让AI智能体真正听懂人类用自然语言下达的指令,是人工智能领域一个既有实用价值又有科学意义的目标,但当前深度学习方法需要海量数据才能学会,距离让真实人类参与训练还很遥远。为此,研究者构建了一个名为 BabyAI 的研究平台:它包含一个2D网格世界环境和19个难度递进的任务关卡,配套一套组合性强、共有约2.48×10¹⁹种可能表达的合成语言(Baby Language),以及一个能模拟人类教师行为、随时生成示范或给出动作建议的机器人教师。实验结果表明,即便是最简单的关卡,基于神经网络的强化学习和行为克隆(behavioral cloning)智能体也需要数十万乃至数百万次交互才能达到接近完美的表现,充分证明现有深度学习方法在学习具有组合性质的语言时样本效率严重不足。这一发现明确指出了未来研究的核心瓶颈,而 BabyAI 平台本身也作为开源工具发布,旨在推动学界在人机协作式语言学习和样本效率提升方面取得实质进展。
原文 arXiv:1810.08272;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1810.08272v4