The Cost of Training NLP Models A Concise Overview
Or Sharir AI21 Labs \AndBarak Peleg AI21 Labs \AndYoav Shoham AI21 Labs
Abstract
We review the cost of training large-scale language models, and the drivers of these costs. The intended audience includes engineers and scientists budgeting their model-training experiments, as well as non-practitioners trying to make sense of the economics of modern-day Natural Language Processing (NLP).111We thank Barak Lenz, Shai Shalev-Shwartz and other members of AI21 Labs, as well as Jack Clark, Jeff Dean, Deep Ganguli, Chris Re, Sebastian Ruder and Lior Wolf, who generously commented on previous drafts. Further comments on the document are welcome, and the document will be updated as appropriate. Note: While the comments of our colleagues from other organizations greatly improved the document, they were not representing their organizations, did not share any proprietary information, and may not necessarily agree with everything written here.
中文速览
训练一个大规模语言模型到底要花多少钱?这篇综述试图把这件事讲清楚,对象是既要控制预算的工程师和研究者,也包括想搞懂AI经济账的普通读者。文章从实际数据入手,指出训练一个15亿参数的BERT模型单次跑下来就要8万美元,算上调参和多次重跑,总费用可能高达160万美元,而像T5这样的百亿参数模型整个项目的花费可能逼近千万美元量级。驱动成本飙升的核心原因是数据集规模、模型参数量和训练token总量这三类因素同步爆炸式增长,叠加调参所需的大量重复实验,让"隐性成本"远超单次训练本身。尽管硬件降价、更高效的架构(如Reformer、ALBERT)以及行业逐渐收敛的军备竞赛心态有望缓解这一趋势,但文章坦言:只要大模型还能带来性能提升,资金雄厚的大公司就会持续砸钱,这对资源有限的研究者和小型机构来说是一道越来越高的门槛。
原文 arXiv:2004.08900;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2004.08900v1