Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
\fnmXiao \surWang \fnmGuangyao \surChen \fnmGuangwu \surQian \fnmPengcheng \surGao \fnmXiao-Yong \surWei \fnmYaowei \surWang(✉)✉{}^{(\textrm{{\char 0}})}start_FLOATSUPERSCRIPT ( ✉ ) end_FLOATSUPERSCRIPT \fnmYonghong \surTian(✉)✉{}^{(\textrm{{\char 0}})}start_FLOATSUPERSCRIPT ( ✉ ) end_FLOATSUPERSCRIPT \fnmWen \surGao [ [ [ [
Abstract
With the urgent demand for generalized deep models, many pre-trained big models are proposed, such as BERT, ViT, GPT, etc. Inspired by the success of these models in single domains (like computer vision and natural language processing), the multi-modal pre-trained big models have also drawn more and more attention in recent years. In this work, we give a comprehensive survey of these models and hope this paper could provide new insights and helps fresh researchers to track the most cutting-edge works. Specifically, we firstly introduce the background of multi-modal pre-training by reviewing the conventional deep learning, pre-training works in natural language process, computer vision, and speech. Then, we introduce the task definition, key challenges, and advantages of multi-modal pre-training models (MM-PTMs), and discuss the MM-PTMs with a focus on data, objectives, network architectures, and knowledge enhanced pre-training. After that, we introduce the downstream tasks used for the validation of large-scale MM-PTMs, including generative, classification, and regression tasks. We also give visualization and analysis of the model parameters and results on representative downstream
中文速览
多模态预训练大模型(Multi-Modal Pre-Trained Models,MM-PTMs)正在成为人工智能的前沿热点,但现有综述大多只聚焦于视觉与语言两种模态,缺乏系统全面的梳理。这篇论文对该领域进行了迄今较为全面的综述,从数据收集与清洗、网络架构设计、预训练目标设定,到知识增强预训练,逐一梳理了各类 MM-PTMs 的核心组成要素,同时覆盖了音频、视频、表格等更多模态。作者系统整理了用于评估这些模型的下游任务(包括生成、分类、回归等),并对主流模型的参数规模与性能表现进行了横向对比分析。这项工作为初入该领域的研究者提供了一条从历史脉络到最新进展的快速通道,也为后续在多模态大模型上的研究指出了若干值得深入探索的方向。
原文 arXiv:2302.10035;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2302.10035v3