Pretrained Language Models for Text Generation: A Survey
Junyi Li1,3111Equal contribution. Tianyi Tang Wayne Xin Zhao1,3222Corresponding author.、Ji-Rong Wen1,2,3 1Gaoling School of Artificial Intelligence, Renmin University of China 2School of Information, Renmin University of China 3Beijing Key Laboratory of Big Data Management and Analysis Methods
Abstract
Text generation has become one of the most important yet challenging tasks in natural language processing (NLP). The resurgence of deep learning has greatly advanced this field by neural generation models, especially the paradigm of pretrained language models (PLMs). In this paper, we present an overview of the major advances achieved in the topic of PLMs for text generation. As the preliminaries, we present the general task definition and briefly describe the mainstream architectures of PLMs for text generation. As the core content, we discuss how to adapt existing PLMs to model different input data and satisfy special properties in the generated text. We further summarize several important fine-tuning strategies for text generation. Finally, we present several future directions and conclude this paper. Our survey aims to provide text generation researchers a synthesis and pointer to related research.
中文速览
预训练语言模型(pretrained language models, PLMs)如BERT、GPT等已经深刻改变了自然语言处理,但如何系统地将它们用于文本生成这一核心任务,此前缺乏全面的综述。这篇论文从任务定义、模型架构出发,系统梳理了PLMs在文本生成中的主要进展:一方面讨论如何让PLMs处理不同类型的输入数据(纯文本、知识图谱等结构化数据、图像视频等多媒体数据),另一方面探讨如何让生成文本满足可控性、忠实性、多样性等特殊要求,同时总结了多种针对文本生成的微调策略。研究发现,PLMs凭借在海量语料上积累的语言知识,能显著提升翻译、摘要、对话等各类生成任务的效果,但在长文档建模、多模态融合、低资源场景等方面仍面临挑战。这项工作是首个聚焦PLMs与文本生成交叉领域的系统性综述,为该领域研究者提供了清晰的研究地图和方向指引。
原文 arXiv:2105.10311;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2105.10311v2