A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT
Ce Zhou Note: The authors contributed equally to this research. Correspondence to Ce and Qian Li Qian Li Chen Li Jun Yu Yixin Liu Guangjing Wang Kai Zhang Affiliation: Michigan State University, Beihang University, Lehigh University, Cheng Ji Qiben Yan Lifang He Affiliation: Michigan State University, Beihang University, Lehigh University, Hao Peng Jianxin Li Jia Wu Ziwei Liu Pengtao Xie Affiliation: Macquarie University, Nanyang Technological University, University of California San Diego, Caiming Xiong Jian Pei Philip S. Yu Affiliation: Salesforce AI Research,Duke University, University of Illinois at Chicago Lichao Sun Affiliation: Michigan State University, Beihang University, Lehigh University,
Abstract
Pretrained Foundation Models (PFMs) are regarded as the foundation for various downstream tasks with different data modalities. A PFM (e.g., BERT, ChatGPT, and GPT-4) is trained on large-scale data which provides a reasonable parameter initialization for a wide range of downstream applications. In contrast to earlier approaches that utilize convolution and recurrent modules to extract features, BERT learns bidirectional encoder representations from Transformers, which are trained on large datasets as contextual language models. Similarly, the Generative Pretrained Transformer (GPT) method employs Transformers as the feature extractor and is trained using an autoregressive paradigm on large datasets. Recently, ChatGPT shows promising success on large language models, which applies an autoregressive language model with zero shot or few shot prompting. The remarkable achievements of PFM have brought significant breakthroughs to various fields of AI in recent years. Numerous studies have proposed different methods, datasets, and evaluation metrics, raising the demand for an updated survey.
原文 arXiv:2302.09419;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2302.09419v3