Achieving Human Parity on Automatic Chinese to English News Translation
Hany Hassan111Corresponding author: Microsoft AI、Research Anthony Aue Microsoft AI、Research Chang Chen Microsoft AI、Research Vishal Chowdhary Microsoft AI、Research Jonathan Clark Microsoft AI、Research Christian Federmann Microsoft AI、Research Xuedong Huang Microsoft AI、Research Marcin Junczys-Dowmunt Microsoft AI、Research William Lewis Microsoft AI、Research Mu Li Microsoft AI、Research Shujie Liu Microsoft AI、Research Tie-Yan Liu Microsoft AI、Research Renqian Luo Microsoft AI、Research Arul Menezes Microsoft AI、Research Tao Qin Microsoft AI、Research Frank Seide Microsoft AI、Research Xu Tan Microsoft AI、Research Fei Tian Microsoft AI、Research Lijun Wu Microsoft AI、Research Shuangzhi Wu Microsoft AI、Research Yingce Xia Microsoft AI、Research Dongdong Zhang Microsoft AI、Research Zhirui Zhang Microsoft AI、Research Ming Zhou Microsoft AI、Research
Abstract
Machine translation has made rapid advances in recent years. Millions of people are using it today in online translation systems and mobile applications in order to communicate across language barriers. The question naturally arises whether such systems can approach or achieve parity with human translations. In this paper, we first address the problem of how to define and accurately measure human parity in translation. We then describe Microsoft’s machine translation system and measure the quality of its translations on the widely used WMT 2017 news translation task from Chinese to English. We find that our latest neural machine translation system has reached a new state-of-the-art, and that the translation quality is at human parity when compared to professional human translations. We also find that it significantly exceeds the quality of crowd-sourced non-professional translations.
中文速览
微软研究团队挑战"机器翻译能否媲美人类"这一核心难题,以中英新闻翻译为测试场景,在广泛使用的WMT 2017评测基准上系统评估其最新神经机器翻译(Neural Machine Translation, NMT)系统的表现。为突破现有NMT的瓶颈,他们综合运用了多项技术:利用翻译任务天然的双向对称性进行联合训练(Dual Learning)、引入二次解码精修机制(Deliberation Networks)、采用左右双向解码一致性约束降低逐步生成时的误差积累,以及严格的数据筛选与系统融合(system combination)。在人工评估上,他们放弃了容易受参考译文质量影响的自动指标,转而采用基于源文本的直接打分(source-based direct assessment)并进行统计显著性检验,以此严格定义"人类水准"。最终结果显示,该系统的译文质量与专业人工翻译在统计上无显著差异,达到了人类水准,且明显优于众包非专业翻译——这是机器翻译领域首次以严格统计方法证实这一里程碑,对推动翻译技术实用化具有重要意义。
原文 arXiv:1803.05567;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1803.05567v2