Attention Is All You Need
Ashish Vaswani Google Brain、Noam Shazeer Google Brain、Niki Parmar Google Research、Jakob Uszkoreit Google Research、Llion Jones Google Research、Aidan N. Gomez University of Toronto、Łukasz Kaiser Google Brain、Illia Polosukhin Equal contribution. Listing order is random. Jakob proposed replacing RNNs with self-attention and started the effort to evaluate this idea. Ashish, with Illia, designed and implemented the first Transformer models and has been crucially involved in every aspect of this work. Noam proposed scaled dot-product attention, multi-head attention and the parameter-free position representation and became the other person involved in nearly every detail. Niki designed, implemented, tuned and evaluated countless model variants in our original codebase and tensor2tensor. Llion also experimented with novel model variants, was responsible for our initial codebase, and efficient inference and visualizations. Lukasz and Aidan spent countless long days designing various parts of and implementing tensor2tensor, replacing our earlier codebase, greatly improving results and massively accelerating our research. Work performed while at Google Brain.Work performed while at Google Research.
Abstract
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.
中文速览
机器翻译长期依赖循环神经网络(RNN)处理序列,但这类模型必须逐步计算、无法并行,训练效率低下。Transformer 彻底抛弃了循环结构和卷积,完全基于"注意力机制"(attention mechanism)来捕捉序列中任意位置之间的依赖关系,编码器和解码器均由多层自注意力与前馈网络堆叠而成。在机器翻译标准测试集上,Transformer 在英德翻译中获得 28.4 的 BLEU 分数、在英法翻译中达到 41.8,均超越此前所有模型,且仅用 3.5 天、8 块 GPU 就完成训练,训练成本大幅低于同类最佳方案。这项工作开创了以纯注意力架构统一序列建模的新范式,为后来 BERT、GPT 等大语言模型奠定了基础。
原文 arXiv:1706.03762;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1706.03762v7