Multimodal Foundation Models: From Specialists to General-Purpose Assistants
Chunyuan Li∗♠, Zhe Gan∗, Zhengyuan Yang∗, Jianwei Yang∗, Linjie Li∗, Lijuan Wang, Jianfeng Gao Microsoft Corporation ∗ Core Contribution ♠ Project Lead
Abstract
This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to general-purpose assistants. The research landscape encompasses five core topics, categorized into two classes. $(i)$ We start with a survey of well-established research areas: multimodal foundation models pre-trained for specific purposes, including two topics – methods of learning vision backbones for visual understanding and text-to-image generation. $(ii)$ Then, we present recent advances in exploratory, open research areas: multimodal foundation models that aim to play the role of general-purpose assistants, including three topics – unified vision models inspired by large language models (LLMs), end-to-end training of multimodal LLMs, and chaining multimodal tools with LLMs. The target audiences of the paper are researchers, graduate students, and professionals in computer vision and vision-language multimodal communities who are eager to learn the basics and recent advances in multimodal foundation models.
中文速览
视觉智能正从"专才"走向"通才"——这篇综述系统梳理了多模态基础模型(multimodal foundation models)在视觉和视觉-语言领域的发展脉络,涵盖视觉理解、文本驱动图像生成、统一视觉模型、多模态大语言模型(multimodal LLMs)以及用大语言模型串联多模态工具五大核心主题。作者将这些工作分为两类:一类是已相对成熟的专用预训练模型,另一类是受大语言模型启发、正在兴起的通用视觉助手研究;通过梳理从CLIP、DALL-E到LLaVA、GPT-4V等代表性工作,勾勒出一条从任务专用到通用智能助手的清晰演进路径。这篇综述的重要意义在于,它不只是文献整理,更提出了一个明确视角:就像ChatGPT统一了语言任务,构建能理解和生成视觉内容、遵从人类意图的通用视觉助手,正成为计算机视觉领域下一阶段最关键的研究方向。
原文 arXiv:2309.10020;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2309.10020v1