UL2: Unifying Language Learning Paradigms
Yi Tay Mostafa Dehghani Thanks: Yi and Mostafa are co-leads of this project and are denoted with $ˆ*$. $♯$ denotes technical research contributors. $♭$ denotes data、infrastructure contributions. $ˆ△$ denotes advising contributions. Don, denoted with $ˆ□$ is the last author. Full contributions of all authors at the end of paper. Correspondence to or Vinh Q. Tran Xavier Garcia Jason Wei Xuezhi Wang Hyung Won Chung Siamak Shakeri Dara Bahri Tal Schuster Huaixiu Steven Zheng Denny Zhou Neil Houlsby Donald Metzler Google Brain
Abstract
Existing pre-trained models are generally geared towards a particular class of problems. To date, there seems to be still no consensus on what the right architecture and pre-training setup should be. This paper presents a unified framework for pre-training models that are universally effective across datasets and setups. We begin by disentangling architectural archetypes with pre-training objectives – two concepts that are commonly conflated. Next, we present a generalized and unified perspective for self-supervision in NLP and show how different pre-training objectives can be cast as one another and how interpolating between different objectives can be effective. We then propose Mixture-of-Denoisers (MoD), a pre-training objective that combines diverse pre-training paradigms together. We furthermore introduce a notion of mode switching, wherein downstream fine-tuning is associated with specific pre-training schemes. We conduct extensive ablative experiments to compare multiple pre-training objectives and find that our method pushes the Pareto-frontier by outperforming T5 and/or GPT-like models across multiple diverse setups. Finally, by scaling our model up to 20B parameters, we a
原文 arXiv:2205.05131;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.05131v3