Agent AI: Surveying the Horizons of Multimodal Interaction
Zane Durante1†, Qiuyuan Huang2‡∗, Naoki Wake2∗, Ran Gong3†, Jae Sung Park4†, Bidipta Sarkar1†, Rohan Taori1†, Yusuke Noda5, Demetri Terzopoulos3, Yejin Choi4, Katsushi Ikeuchi2, Hoi Vo5, Li Fei-Fei1, Jianfeng Gao2 1Stanford University; 2Microsoft Research, Redmond; 3University of California, Los Angeles; 4University of Washington; 5Microsoft Gaming Equal Contribution. ‡ Project Lead. † Work done while interning at Microsoft Research, Redmond.
Abstract
Multi-modal AI systems will likely become a ubiquitous presence in our everyday lives. A promising approach to making these systems more interactive is to embody them as agents within physical and virtual environments. At present, systems leverage existing foundation models as the basic building blocks for the creation of embodied agents. Embedding agents within such environments facilitates the ability of models to process and interpret visual and contextual data, which is critical for the creation of more sophisticated and context-aware AI systems. For example, a system that can perceive user actions, human behavior, environmental objects, audio expressions, and the collective sentiment of a scene can be used to inform and direct agent responses within the given environment. To accelerate research on agent-based multimodal intelligence, we define “Agent AI” as a class of interactive systems that can perceive visual stimuli, language inputs, and other environmentally-grounded data, and can produce meaningful embodied actions. In particular, we explore systems that aim to improve agents based on next-embodied action prediction by incorporating external knowledge, multi-sensory inpu
中文速览
多模态AI智能体(Multimodal Agent AI)如何从各自孤立的感知、规划、执行模块走向真正整合的通用智能,是当前AI研究的核心挑战。这篇综述提出"Agent AI"这一概念框架,将能够感知视觉、语言及其他环境信号、并能在物理或虚拟环境中产生有意义行动的交互系统统一归纳其中,重点探讨如何借助大语言模型(LLM)和视觉语言模型(VLM)、结合外部知识、多感官输入与人类反馈来提升智能体的下一步动作预测能力。文章系统梳理了Agent AI在游戏(含VR/AR)、机器人和医疗健康三大应用领域的方法论与评测基准,并介绍了作者团队专为多模态智能体训练构建的新数据集。研究表明,将AI系统嵌入有实体交互的环境中,不仅能驱动更复杂的情境感知,还能有效缓解大型基础模型生成与现实环境不符的"幻觉"问题;这一方向对于推动AI从被动的结构化任务走向动态、自主的通用智能体具有重要意义。
原文 arXiv:2401.03568;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2401.03568v2