Do Transformer Modifications Transfer Across Implementations and Applications?
Sharan Narang、Hyung Won Chung、Yi Tay、William Fedus \ANDThibault Fevry、Michael Matena、Karishma Malkan、Noah Fiedel \ANDNoam Shazeer、Zhenzhong Lan、Yanqi Zhou、Wei Li \ANDNan Ding、Jake Marcus、Adam Roberts、Colin Raffel Correspondence to Work completed while at Google
Abstract
The research community has proposed copious modifications to the Transformer architecture since it was introduced over three years ago, relatively few of which have seen widespread adoption. In this paper, we comprehensively evaluate many of these modifications in a shared experimental setting that covers most of the common uses of the Transformer in natural language processing. Surprisingly, we find that most modifications do not meaningfully improve performance. Furthermore, most of the Transformer variants we found beneficial were either developed in the same codebase that we used or are relatively minor changes. We conjecture that performance improvements may strongly depend on implementation details and correspondingly make some recommendations for improving the generality of experimental results.
中文速览
三年多来研究者提出了大量针对Transformer架构的改进方案,但真正被广泛采用的寥寥无几,这篇论文试图弄清背后的原因。作者在统一的实验框架下,横跨迁移学习与机器翻译两大场景,系统测试了激活函数、归一化方式、参数共享、注意力变体等数十种改进方案。结果令人意外:绝大多数改进在统一测试环境下并不能带来显著的性能提升,而那些确实有效的变体,要么改动极其微小,要么本就是在同一套代码库中开发出来的。这一发现揭示了一个不容忽视的问题——Transformer改进的效果很可能高度依赖具体的实现细节和实验环境,难以跨代码库、跨任务泛化,提示社区在评估和发表架构改进时需要更严格、更多样化的实验设计。
原文 arXiv:2102.11972;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2102.11972v2