Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
\nameAditya Mogadala \nameMarimuthu Kalimuthu \nameDietrich Klakow \addrSpoken Language Systems (LSV) Saarland Informatics Campus Saarland University 66123 Saarbrücken, Germany
Abstract
Interest in Artificial Intelligence (AI) and its applications has seen unprecedented growth in the last few years. This success can be partly attributed to the advancements made in the sub-fields of AI such as machine learning, computer vision, and natural language processing. Much of the growth in these fields has been made possible with deep learning, a sub-area of machine learning that uses artificial neural networks. This has created significant interest in the integration of vision and language. In this survey, we focus on ten prominent tasks that integrate language and vision by discussing their problem formulation, methods, existing datasets, evaluation measures, and compare the results obtained with corresponding state-of-the-art methods. Our efforts go beyond earlier surveys which are either task-specific or concentrate only on one type of visual content, i.e., image or video. Furthermore, we also provide some potential future directions in this field of research with an anticipation that this survey stimulates innovative thoughts and ideas to address the existing challenges and build new applications.
中文速览
视觉与语言的深度融合正成为人工智能最活跃的前沿方向,但此前的综述要么只聚焦单一任务,要么只讨论图像或视频中的某一类视觉内容。这篇综述系统梳理了十项将语言与视觉结合的核心任务,涵盖视觉描述生成、视觉问答、视觉对话、视觉推理、视觉故事生成、视觉指代表达、视觉蕴含、多模态机器翻译、视觉生成以及视觉导航,对每项任务的问题定义、常用方法、现有数据集、评估指标和最新结果均作了深入介绍,并专门讨论了近年兴起的视觉-语言联合预训练(joint vision-language pretraining)范式。研究发现,深度学习尤其是Transformer架构的突破极大推动了各任务性能的提升,但跨模态理解、常识推理、数据偏差等挑战仍未解决。这篇综述为希望进入该领域的研究者提供了一站式参考,也为未来开发更智能的人机交互、辅助视障人士、自动驾驶等实际应用指明了方向。
原文 arXiv:1907.09358;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1907.09358v3