Aha.
正在载入中英对照阅读…

arXiv:1810.08272 · 中英对照阅读

BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning

Maxime Chevalier-Boisvert、Dzmitry Bahdanau、Salem Lahlou、Lucas Willems、Chitwan Saharia、Thien Huu Nguyen、Yoshua Bengio

中文速览

让人类通过语言指令教会人工智能在真实环境中行动,最大的难题是如何用尽量少的人类示范和反馈学会组合式语言。研究者搭建了 BabyAI 平台:在可部分观察、能开门和搬动物体的二维网格世界中,设计了由易到难的 19 个任务和一套可组合的合成语言,并用能随时示范和指导的机器人模拟人类教师,再测试行为克隆、强化学习、预训练和交互式模仿学习等方法。实验显示,现有深度学习方法要把这些任务学好仍需大量示范或交互,面对具有组合结构的语言时样本效率明显不够,预训练和互动教学虽有帮助但尚未根本解决问题。这个平台因此既提供了统一的研究环境和基准,也明确指出了让真实人类高效教会智能体理解语言前最需要突破的瓶颈。

摘要

Allowing humans to interactively train artificial agents to understand language instructions is desirable for both practical and scientific reasons. Though, given the lack of sample efficiency in current learning methods, reaching this goal may require substantial research efforts. We introduce the BabyAI research platform, with the goal of supporting investigations towards including humans in the loop for grounded language learning. The BabyAI platform comprises an extensible suite of 19 levels of increasing difficulty. Each level gradually leads the agent towards acquiring a combinatorially rich synthetic language, which is a proper subset of English. The platform also provides a hand-crafted bot agent, which simulates a human teacher. We report estimated amount of supervision required for training neural reinforcement and behavioral-cloning agents on some BabyAI levels. We put forward strong evidence that current deep learning methods are not yet sufficiently sample-efficient in the context of learning a language with compositional properties.

术语表

BabyAI
BabyAI
MiniGrid
MiniGrid
grounded language learning
具身语言学习
human-in-the-loop
人在回路中
sample efficiency
样本效率
curriculum learning
课程学习
interactive teaching
交互式教学
gridworld
网格世界
synthetic language
合成语言
Baby Language
Baby语言
compositional language
组合式语言
context-free grammar
上下文无关文法
partial observability
部分可观测性
imitation learning
模仿学习
behavioral cloning
行为克隆
reinforcement learning
强化学习
deep learning
深度学习
interactive imitation learning
交互式模仿学习
DAGGER
DAGGER
TAMER
TAMER
Searn
Searn
maximum-entropy RL
最大熵强化学习
bAbI tasks
bAbI任务
SAIL
SAIL
Room-to-Room
Room-to-Room