Transcending Scaling Laws with 0.1% Extra Compute
Yi Tay Jason Wei Hyung Won Chung Vinh Q. Tran David R. So Siamak Shakeri Xavier Garcia Huaixiu Steven Zheng Jinfeng Rao Aakanksha Chowdhery Denny Zhou Donald Metzler Slav Petrov Neil Houlsby Quoc V. Le Mostafa Dehghani Google
Abstract
Scaling language models improves performance but comes with significant computational costs. This paper proposes UL2R, a method that substantially improves existing language models and their scaling curves with a relatively tiny amount of extra compute. The key idea is to continue training a state-of-the-art large language model (e.g., PaLM) on a few more steps with UL2’s mixture-of-denoiser objective. We show that, with almost negligible extra computational costs and no new sources of data, we are able to substantially improve the scaling properties of large language models on downstream metrics. In this paper, we continue training PaLM with UL2R, introducing a new set of models at 8B, 62B, and 540B scale which we call U-PaLM. Impressively, at 540B scale, we show an approximately 2x computational savings rate where U-PaLM achieves the same performance as the final PaLM 540B model at around half its computational budget (i.e., saving $\sim$ 4.4 million TPUv4 hours).
原文 arXiv:2210.11399;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2210.11399v2