MT-Adapted Datasheets for Datasets: Template and Repository
Marta R. Costa-jussà, Roger Creus, Oriol Domingo, Albert Domínguez, Miquel Escobar, Cayetana López, Marina Garcia and Margarita Geleta TALP Research Center Universitat Politècnica de Catalunya, Barcelona
Abstract
In this report we are taking the standardized model proposed by Gebru et al. [Gebru et al., 2018] for documenting the popular machine translation datasets of the EuroParl [Koehn, 2005] and News-Commentary [Barrault et al., 2019]. Within this documentation process, we have adapted the original datasheet to the particular case of data consumers within the Machine Translation area. We are also proposing a repository for collecting the adapted datasheets in this research area.
中文速览
机器翻译(Machine Translation, MT)领域长期依赖 EuroParl、News-Commentary 等大型语料库,但这些数据集几乎没有任何标准化的文档记录,导致研究者对其偏见、来源和使用限制知之甚少。为此,作者以 Gebru 等人提出的"数据集说明书"(Datasheets for Datasets)框架为基础,针对机器翻译数据消费者的实际需求对其进行了适配——删去了创建者才能回答的问题,并补充了语言覆盖均衡性、性别与社会偏见、文本正式程度、合规性保障等机器翻译场景下更为关键的问题。以修改后的模板为示范,他们实际填写了 EuroParl 和 News-Commentary 两份数据集说明书,并搭建了一个开放的在线仓库供社区持续贡献。这项工作的意义在于:它为整个机器翻译社区提供了一套轻量可行的透明度工具,有助于在使用数据之前就发现并正视潜在的数据偏见,从而推动更公平、更负责任的自然语言处理研究。
原文 arXiv:2005.13156;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2005.13156v1