Beyond English-Centric Multilingual Machine Translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli†, Armand Joulin Facebook AI, ∗LORIA Corresponding Author. Email: contribution. Order determined with coin flip.Holger Schwenk, Onur Celebi, Vishrav Chaudhary, Ahmed El-Kishky, Angela Fan, Edouard Grave, Zhiyi Ma, and Guillaume Wenzek worked on large-scale data mining, including improvements to fasttext, LASER, CCAligned, and CCMatrix. Siddharth Goyal worked on backtranslation. Michael Auli, Mandeep Baines, Shruti Bhosale, Tom Birch, Sergey Edunov, Angela Fan, Naman Goyal, Siddharth Goyal, and Vitaliy Liptchinsky worked on model scaling and scaling infrastructure. Michael Auli, Shruti Bhosale, Sergey Edunov, Angela Fan, Edouard Grave, Armand Joulin, Zhiyi Ma, and Holger Schwenk worked on model and experimental design. Michael Auli, Ahmed El-Kishky, Angela Fan, Edouard Grave, Armand Joulin, Zhiyi Ma, and Holger Schwenk wrote the paper.
Abstract
Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only on data which was translated from or to English. While this is supported by large sources of training data, it does not reflect translation needs worldwide. In this work, we create a true Many-to-Many multilingual translation model that can translate directly between any pair of 100 languages. We build and open source a training dataset that covers thousands of language directions with supervised data, created through large-scale mining. Then, we explore how to effectively increase model capacity through a combination of dense scaling and language-specific sparse parameters to create high quality models. Our focus on non-English-Centric models brings gains of more than 10 BLEU when directly translating between non-English directions while performing competitively to the best single systems of WMT. We open-source our scripts so that others may reproduce the data, evaluation, and final M2M-100 model here.
中文速览
多语言机器翻译长期以来以英语为中心,只训练"某语言↔英语"的方向,导致非英语语言对之间的直接翻译质量很差。为解决这一问题,研究团队通过大规模自动挖掘构建了覆盖100种语言、75亿句对的真正"多对多"平行语料库,让模型能直接在任意语言对之间翻译,而无需经英语中转。在模型架构上,他们结合稠密扩容与语言专属的稀疏混合专家(Mixture-of-Experts)机制,将模型参数量扩展至154亿,训练得到M2M-100模型。实验结果显示,与以英语为中心的方案相比,该模型在非英语翻译方向上BLEU提升超过10分,同时在WMT等权威评测中也能与最优单系统媲美。这项工作首次从数据到模型系统性地打破了多语言翻译对英语的依赖,相关数据、代码和模型均已开源,对推动全球语言平等覆盖具有重要意义。
原文 arXiv:2010.11125;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2010.11125v1