Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu Soricut Google Research
Abstract
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets. In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically-diverse set of $3600$ images annotated with human-generated reference captions in $36$ languages. The images were selected from across the world, covering regions where the $36$ languages are spoken, and annotated with captions that achieve consistency in terms of style across all languages, while avoiding annotation artifacts due to direct translation. We apply this benchmark to model selection for massively multilingual image captioning models, and show strong correlation results with human evaluations when using XM3600 as golden references for automatic metrics.
中文速览
多语言图像描述(image captioning)领域长期缺乏高质量的评测数据集,导致模型好坏只能靠昂贵的人工评估来判断。为此,研究者构建了 Crossmodal-3600(XM3600)数据集:从全球各地精心挑选3600张具有地理多样性的图片,并由双语专业标注员为36种语言各自独立撰写描述文字,而非依赖机器翻译——标注员先浏览英文自动生成的描述以统一风格,再在不看原文的情况下用目标语言重新描写图片,从而避免了翻译腔(translation artifacts)的干扰。实验表明,用XM3600作为参考答案计算CIDEr等自动指标所得到的模型排名,与人工评估结果高度吻合。这一数据集的发布为大规模多语言视觉语言研究提供了可靠、可复现的评测基准,大幅降低了未来模型比较对人工评估的依赖。
原文 arXiv:2205.12522;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.12522v2