Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent
Zhilin Yang, Saizheng Zhang, Jack Urbanek, Will Feng, Alexander H. Miller Arthur Szlam, Douwe Kiela、Jason Weston Facebook AI Research
Abstract
Contrary to most natural language processing research, which makes use of static datasets, humans learn language interactively, grounded in an environment. In this work we propose an interactive learning procedure called Mechanical Turker Descent (MTD) and use it to train agents to execute natural language commands grounded in a fantasy text adventure game. In MTD, Turkers compete to train better agents in the short term, and collaborate by sharing their agents’ skills in the long term. This results in a gamified, engaging experience for the Turkers and a better quality teaching signal for the agents compared to static datasets, as the Turkers naturally adapt the training data to the agent’s abilities.
中文速览
现有NLP研究大多依赖静态数据集,而人类其实是在与环境互动中学习语言的——为弥合这一差距,研究者提出了一种名为"机械土耳其人下降"(Mechanical Turker Descent,MTD)的交互式数据收集框架:让众包标注者在一款名为"征服地下城"的奇幻文字冒险游戏中,通过多轮竞争与协作来训练语言智能体,每轮内标注者各自为自己的智能体提供训练样例并相互竞争排名奖金,每轮结束后所有人的数据合并共享,使智能体在长期协作中持续进步。实验表明,无论是标准的Seq2Seq模型还是专为图结构环境设计的AC-Seq2Seq模型,用MTD训练的效果都显著优于静态数据集基线,消融研究进一步证实"对标注者有吸引力"和"让训练数据难度与智能体能力相匹配"是MTD成功的两大关键。这项工作为构建能真正理解并执行自然语言指令的具身智能体提供了一条可扩展、可复用的人机协同训练路径。
原文 arXiv:1711.07950;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1711.07950v3