HoME: a Household Multimodal Environment
Simon Brodeur1, Ethan Perez2,3, Ankesh Anand2††, Florian Golemo2,4††, Luca Celotti1, Florian Strub2,5, Jean Rouat1, Hugo Larochelle6,7, Aaron Courville2,7 1Université de Sherbrooke, 2MILA, Université de Montréal, 3Rice University, 4INRIA Bordeaux, 5Univ. Lille, Inria, UMR 9189 - CRIStAL, 6Google Brain, 7CIFAR Fellow {simon.brodeur, luca.celotti, {florian.golemo, {ankesh.anand, These authors contributed equally.
Abstract
We introduce HoME: a Household Multimodal Environment for artificial agents to learn from vision, audio, semantics, physics, and interaction with objects and other agents, all within a realistic context. HoME integrates over 45,000 diverse 3D house layouts based on the SUNCG dataset, a scale which may facilitate learning, generalization, and transfer. HoME is an open-source, OpenAI Gym-compatible platform extensible to tasks in reinforcement learning, language grounding, sound-based navigation, robotics, multi-agent learning, and more. We hope HoME better enables artificial agents to learn as humans do: in an interactive, multimodal, and richly contextualized setting.
中文速览
让AI智能体像人类一样在真实家庭环境中通过视觉、听觉、语言和肢体互动来学习,长期以来缺乏一个足够大规模、足够多模态的训练平台,这正是HoME(家庭多模态环境,Household Multimodal Environment)要填补的空白。研究团队基于SUNCG数据集构建了包含超过45,000套手工设计3D房屋的开放平台,集成了视觉渲染、物理声学光线追踪、语义标注、物理引擎和多智能体支持,智能体可以在其中自由移动、拾取物品、倾听环境声音并与其他智能体交互。实验验证该平台可在单核CPU上实现超实时运行,并支持GPU加速和并行实例,在此基础上可扩展出指令跟随、视觉问答、声源定位等多种任务。HoME的重要性在于,它是首个同时具备高保真音频、真实物理模拟和大规模多样场景的开源交互平台,为推动具身通用人工智能的研究提供了迄今最完整的多模态训练基础设施。
原文 arXiv:1711.11017;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1711.11017v1