M6: A Chinese Multimodal Pretrainer
Junyang Lin1∗, Rui Men1∗, An Yang1∗, Chang Zhou1, Ming Ding2, Yichang Zhang1, Peng Wang1, Ang Wang1, Le Jiang1, Xianyan Jia1, Jie Zhang1, Jianwei Zhang1, Xu Zou2, Zhikang Li1, Xiaodong Deng1, Jie Liu1, Jinbao Xue1, Huiling Zhou1, Jianxin Ma1, Jin Yu1, Yong Li1, Wei Lin1, Jingren Zhou1, Jie Tang2†, Hongxia Yang1† 1Alibaba GroupChina 2Tsinghua UniversityChina junyang.ljy, menrui.mr, ya235025, ericzhou.zc, yichang.zyc, wangang.wa, jiangle.jl, xianyan.xianyanjia, wanglin.zj, zhikang.lzk, xiaodongdeng.dxd, sanshuai.lj, zhiji.xjb, zhule.zhl, jason.mjx, jiufeng.ly, weilin.lw, jingren.zhou, dm18,
Abstract
In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains. We propose a cross-modal pretraining method called M6, referring to Multi-Modality to Multi-Modality Multitask Mega-transformer, for unified pretraining on the data of single modality and multiple modalities. We scale the model size up to 10 billion and 100 billion parameters, and build the largest pretrained model in Chinese. We apply the model to a series of downstream applications, and demonstrate its outstanding performance in comparison with strong baselines. Furthermore, we specifically design a downstream task of text-guided image generation, and show that the finetuned M6 can create high-quality images with high resolution and abundant details.
中文速览
中文互联网上长期缺乏用于多模态预训练的大规模高质量数据集和对应的超大模型,为此研究团队构建了迄今最大的中文多模态语料库(超过1.9TB图像、292GB文本),并在此基础上提出了名为M6(Multi-Modality to Multi-Modality Multitask Mega-transformer)的跨模态预训练框架。M6采用统一的编解码器架构,通过文本去噪、图像到文本生成、多模态到文本生成等多任务联合预训练,使模型同时具备单模态理解与跨模态生成能力,并将参数规模扩展到100亿乃至1000亿,成为当时最大的中文预训练模型。在视觉问答、图文匹配、图像描述生成等多个下游任务上,M6均大幅超越强基线,还首次将预训练能力延伸到文本生成高分辨率图像的任务中。这项工作填补了中文多模态大模型的空白,也为工业界大规模部署多模态预训练提供了可行的技术路径。
原文 arXiv:2103.00823;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2103.00823v4