Depth Growing for Neural Machine Translation
Lijun Wu Affiliation: School of Data and Computer Science, Sun Yat-sen University; Yiren Wang Thanks: The first two authors contributed equally to this work. This work is conducted at Microsoft Research Asia. Affiliation: {wulijun3, Yingce Xia Thanks: Corresponding author. Affiliation: University of Illinois at Urbana-Champaign; Microsoft Research Asia; Affiliation: {Yingce.Xia, fetia, feiga, taoqin, Fei Tian Affiliation: University of Illinois at Urbana-Champaign; Microsoft Research Asia; Affiliation: {Yingce.Xia, fetia, feiga, taoqin, Fei Gao Affiliation: University of Illinois at Urbana-Champaign; Microsoft Research Asia; Affiliation: {Yingce.Xia, fetia, feiga, taoqin, Tao Qin, Jianhuang Lai, Tie-Yan Liu Affiliation: School of Data and Computer Science, Sun Yat-sen University; Affiliation: University of Illinois at Urbana-Champaign; Microsoft Research Asia; Affiliation: University of Illinois at Urbana-Champaign; Microsoft Research Asia; Affiliation: {Yingce.Xia, fetia, feiga, taoqin, Affiliation: {Yingce.Xia, fetia, feiga, taoqin,
Abstract
While very deep neural networks have shown effectiveness for computer vision and text classification applications, how to increase the network depth of neural machine translation (NMT) models for better translation quality remains a challenging problem. Directly stacking more blocks to the NMT model results in no improvement and even reduces performance. In this work, we propose an effective two-stage approach with three specially designed components to construct deeper NMT models, which results in significant improvements over the strong Transformer baselines on WMT $14$ English $\to$ German and English $\to$ French translation tasks11 1 Our code is available at https://github.com/apeterswu/Depth_Growing_NMT.
原文 arXiv:1907.01968;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1907.01968v1