M6: A Chinese Multimodal Pretrainer
Junyang Lin1∗, Rui Men1∗, An Yang1∗, Chang Zhou1, Ming Ding2, Yichang Zhang1, Peng Wang1, Ang Wang1, Le Jiang1, Xianyan Jia1, Jie Zhang1, Jianwei Zhang1, Xu Zou2, Zhikang Li1, Xiaodong Deng1, Jie Liu1, Jinbao Xue1, Huiling Zhou1, Jianxin Ma1, Jin Yu1, Yong Li1, Wei Lin1, Jingren Zhou1, Jie Tang2†, Hongxia Yang1† Affiliation: 1Alibaba GroupChina 2Tsinghua University, China email: junyang.ljy, menrui.mr, ya235025, ericzhou.zc, yichang.zyc, email: wangang.wa, jiangle.jl, xianyan.xianyanjia, wanglin.zj, email: zhikang.lzk, xiaodongdeng.dxd, sanshuai.lj, zhiji.xjb, zhule.zhl, jason.mjx, email: jiufeng.ly, weilin.lw, jingren.zhou, email: dm18,
Abstract
In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains. We propose a cross-modal pretraining method called M6, referring to Multi-Modality to Multi-Modality Multitask Mega-transformer, for unified pretraining on the data of single modality and multiple modalities. We scale the model size up to 10 billion and 100 billion parameters, and build the largest pretrained model in Chinese. We apply the model to a series of downstream applications, and demonstrate its outstanding performance in comparison with strong baselines. Furthermore, we specifically design a downstream task of text-guided image generation, and show that the finetuned M6 can create high-quality images with high resolution and abundant details.
原文 arXiv:2103.00823;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2103.00823v4