arXiv:1706.03762 · 中英对照阅读
注意力就是你所需要的一切
Attention Is All You Need
中文速览
传统翻译模型按词一个接一个地处理,训练难以并行,也不容易抓住相隔很远的词之间的关系。论文提出 Transformer,用注意力机制让句子里的词彼此“查看”并筛选重要信息,完全去掉循环和卷积结构,同时加入位置信息来保留词序。实验中,它在英德、英法翻译上刷新了当时的成绩,训练速度也更快,还能用于句法分析。它的重要性在于证明:不靠逐步处理序列,也能做出更准、更省时间的语言模型,为后来的大规模语言模型开辟了新路线。
摘要
主流的序列转导模型基于复杂的循环神经网络或卷积神经网络,其中包含一个编码器和一个解码器。性能最佳的模型还通过注意力机制连接编码器和解码器。我们提出了一种全新的简洁网络架构——Transformer,它完全基于注意力机制,彻底摒弃了循环和卷积。在两项机器翻译任务上的实验表明,这些模型在质量上更胜一筹,同时具有更强的并行化能力,训练所需时间也显著更短。在 WMT 2014 英译德翻译任务中,我们的模型取得了 28.4 BLEU 的成绩,较包括集成模型在内的现有最佳结果高出 2 BLEU 以上。在 WMT 2014 英译法翻译任务中,我们的模型创下了单模型 BLEU 评分的新最佳水平,达到 41.8;该模型在八块 GPU 上训练了 3.5 天,训练成本仅为文献中最佳模型的一小部分。我们还通过将 Transformer 成功应用于英语成分句法分析(包括大规模和有限训练数据两种情形),证明了它能够很好地泛化到其他任务。
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.
术语表
- sequence transduction
- 序列转导
- Transformer
- Transformer
- recurrent neural network (RNN)
- 循环神经网络
- convolutional neural network (CNN)
- 卷积神经网络
- encoder-decoder architecture
- 编码器—解码器架构
- attention mechanism
- 注意力机制
- machine translation
- 机器翻译
- BLEU
- BLEU
- WMT 2014
- WMT 2014
- language modeling
- 语言建模
- long short-term memory (LSTM)
- 长短期记忆网络
- gated recurrent neural network
- 门控循环神经网络
- hidden state
- 隐藏状态
- conditional computation
- 条件计算
- self-attention
- 自注意力
- intra-attention
- 内部注意力
- Multi-Head Attention
- 多头注意力
- Extended Neural GPU
- Extended Neural GPU
- ByteNet
- ByteNet
- ConvS2S
- ConvS2S
- end-to-end memory network
- 端到端记忆网络
- auto-regressive
- 自回归
- feed-forward network
- 前馈网络
- residual connection
- 残差连接
- layer normalization
- 层归一化
- Scaled Dot-Product Attention
- 缩放点积注意力
- additive attention
- 加性注意力
- softmax function
- softmax 函数