Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
Xiao Wang Affiliation: Pengcheng Laboratory, Shenzhen 518055, China Affiliation: School of Computer Science and Technology, Anhui University, Hefei 230601, China Guangyao Chen Affiliation: Pengcheng Laboratory, Shenzhen 518055, China Affiliation: School of Computer Science, Peking University, Beijing 100871, China Guangwu Qian Affiliation: Pengcheng Laboratory, Shenzhen 518055, China Pengcheng Gao Affiliation: Pengcheng Laboratory, Shenzhen 518055, China Xiao-Yong Wei Affiliation: Pengcheng Laboratory, Shenzhen 518055, China Affiliation: College of Computer Science, Sichuan University, Chengdu 610065, China Yaowei Wang(\Letter){}^{(\textrm{\Letter})} Affiliation: Pengcheng Laboratory, Shenzhen 518055, China Yonghong Tian(\Letter){}^{(\textrm{\Letter})} Affiliation: Pengcheng Laboratory, Shenzhen 518055, China Affiliation: School of Computer Science, Peking University, Beijing 100871, China Wen Gao Affiliation: Pengcheng Laboratory, Shenzhen 518055, China Affiliation: School of Computer Science, Peking University, Beijing 100871, China
Abstract
With the urgent demand for generalized deep models, many pre-trained big models are proposed, such as BERT, ViT, GPT, etc. Inspired by the success of these models in single domains (like computer vision and natural language processing), the multi-modal pre-trained big models have also drawn more and more attention in recent years. In this work, we give a comprehensive survey of these models and hope this paper could provide new insights and helps fresh researchers to track the most cutting-edge works. Specifically, we firstly introduce the background of multi-modal pre-training by reviewing the conventional deep learning, pre-training works in natural language process, computer vision, and speech. Then, we introduce the task definition, key challenges, and advantages of multi-modal pre-training models (MM-PTMs), and discuss the MM-PTMs with a focus on data, objectives, network architectures, and knowledge enhanced pre-training. After that, we introduce the downstream tasks used for the validation of large-scale MM-PTMs, including generative, classification, and regression tasks. We also give visualization and analysis of the model parameters and results on representative downstream
原文 arXiv:2302.10035;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2302.10035v3