Minimum Risk Training for Neural Machine Translation
Shiqi Shen Affiliation: State Key Laboratory of Intelligent Technology and SystemsTsinghua National Laboratory for Information Science and TechnologyDepartment of Computer Science and Technology, Tsinghua University, Beijing, China Yong Cheng Affiliation: Institute for Interdisciplinary Information Sciences, Tsinghua University, Beijing, China Zhongjun He Affiliation: Baidu Inc., Beijing, China{vicapple22, {hezhongjun, hewei06, {sms, Wei He Affiliation: Baidu Inc., Beijing, China{vicapple22, {hezhongjun, hewei06, {sms, Hua Wu Affiliation: Baidu Inc., Beijing, China{vicapple22, {hezhongjun, hewei06, {sms, Maosong Sun Affiliation: State Key Laboratory of Intelligent Technology and SystemsTsinghua National Laboratory for Information Science and TechnologyDepartment of Computer Science and Technology, Tsinghua University, Beijing, China Yang Liu Thanks: Corresponding author: Yang Liu. Affiliation: State Key Laboratory of Intelligent Technology and SystemsTsinghua National Laboratory for Information Science and TechnologyDepartment of Computer Science and Technology, Tsinghua University, Beijing, China
Abstract
We propose minimum risk training for end-to-end neural machine translation. Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily differentiable. Experiments show that our approach achieves significant improvements over maximum likelihood estimation on a state-of-the-art neural machine translation system across various languages pairs. Transparent to architectures, our approach can be applied to more neural networks and potentially benefit more NLP tasks.
原文 arXiv:1512.02433;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1512.02433v3