Attention Is All You Need
Ashish Vaswani Google Brain、Noam Shazeer Google Brain、Niki Parmar Google Research、Jakob Uszkoreit Google Research、Llion Jones Google Research、Aidan N. Gomez University of Toronto、Łukasz Kaiser Google Brain、Illia Polosukhin Equal contribution. Listing order is random. Jakob proposed replacing RNNs with self-attention and started the effort to evaluate this idea. Ashish, with Illia, designed and implemented the first Transformer models and has been crucially involved in every aspect of this work. Noam proposed scaled dot-product attention, multi-head attention and the parameter-free position representation and became the other person involved in nearly every detail. Niki designed, implemented, tuned and evaluated countless model variants in our original codebase and tensor2tensor. Llion also experimented with novel model variants, was responsible for our initial codebase, and efficient inference and visualizations. Lukasz and Aidan spent countless long days designing various parts of and implementing tensor2tensor, replacing our earlier codebase, greatly improving results and massively accelerating our research. Work performed while at Google Brain.Work performed while at Google Research.
Abstract
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.
中文速览
机器翻译等序列任务长期依赖循环神经网络,计算必须按词逐步进行,训练慢且难以并行。论文提出Transformer,只用注意力机制(attention)和多头自注意力连接序列中各位置,配合位置编码、前馈网络与编码器—解码器结构,完全去掉循环和卷积。实验显示,它在英德翻译上达到28.4 BLEU、英法翻译上达到41.8 BLEU,超过当时最佳系统,同时训练时间和成本大幅降低,并能推广到英文句法分析。其重要性在于,它证明了序列建模不必依赖逐步循环计算,为后来大规模语言模型的发展奠定了核心架构基础。
原文 arXiv:1706.03762;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1706.03762v7