Taiyi-Diffusion-XL: Advancing Bilingual Text-to-Image Generation with Large Vision-Language Model Support
Xiaojun Wu111Equal Contribution. Dixiang Zhang111Equal Contribution. Ruyi Gan♣222Project Leader. Junyu Lu Ziwei Wu Renliang Sun Jiaxing Zhang333Corresponding Author. Pingjian Zhang333Corresponding Author. Yan Song♣333Corresponding Author. International Digital Economy Academy ♣South China University of Technology University of Science and Technology of China {wuxiaojun, zhangdixiang, ganruyi, lujunyu, wuziwei,
Abstract
Recent advancements in text-to-image models have significantly enhanced image generation capabilities, yet a notable gap of open-source models persists in bilingual or Chinese language support. To address this need, we present Taiyi-Diffusion-XL, a new Chinese and English bilingual text-to-image model which is developed by extending the capabilities of CLIP and Stable-Diffusion-XL through a process of bilingual continuous pre-training. This approach includes the efficient expansion of vocabulary by integrating the most frequently used Chinese characters into CLIP’s tokenizer and embedding layers, coupled with an absolute position encoding expansion. Additionally, we enrich text prompts by large vision-language model, leading to better images captions and possess higher visual quality. These enhancements are subsequently applied to downstream text-to-image models. Our empirical results indicate that the developed CLIP model excels in bilingual image-text retrieval. Furthermore, the bilingual image generation capabilities of Taiyi-Diffusion-XL surpass previous models. This research leads to the development and open-sourcing of the Taiyi-Diffusion-XL model, representing a notable adva
中文速览
开源文本生成图像领域长期缺乏能同时理解中英文的高质量模型,Taiyi-Diffusion-XL(太乙-XL)正是为填补这一空白而生。研究团队在已有的CLIP和Stable Diffusion XL基础上进行双语持续预训练:将高频汉字扩充进CLIP的词表与嵌入层、扩展绝对位置编码,并借助大型视觉语言模型Lyrics为训练图片生成更精准的中英文描述,从而同时提升文本编码器和扩散模型的双语理解能力。实验结果显示,改进后的CLIP在中文和英文图文检索基准上均达到最优,太乙-XL在英文COCO和中文COCO-CN数据集上的CLIP相似度、IS和FID三项指标也全面超越Alt-Diffusion、Pai-Diffusion、SD-XL等竞品。这项工作不仅推动了中文图像生成的技术上限,更将模型完整开源,为全球研究者提供了一个真正意义上兼顾中英双语的高质量文生图基础模型。
原文 arXiv:2401.14688;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2401.14688v3