Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
Julia Kreutzer Isaac Caswell Lisa Wang Ahsan Wahab Daan van Esch Nasanbayar Ulzii-Orshikh Affiliation: Google Research, Masakhane NLP, Turkic Interlingua, Haverford College, Allahsera Tapo Nishant Subramani Affiliation: Allen Institute for Artificial Intelligence Artem Sokolov Claytone Sikasote Monang Setyawan Supheakmungkol Sarin Sokhar Samb Affiliation: RobotsMali, Intel Labs, University of Zambia, Google, AIMS-AMMI, Benoît Sagot Clara Rivera Annette Rios Isabel Papadimitriou Affiliation: Inria, University of Zurich, Stanford University, Salomey Osei Affiliation: Kwame Nkrumah University of Science and Technology, Pedro Ortiz Suarez Iroro Orife Kelechi Ogueji Affiliation: Sorbonne Université, Niger-Volta LTI, University of Waterloo Andre Niyongabo Rubungo Toan Q. Nguyen Affiliation: University of Electronic Science and Technology of China, University of Notre Dame, Mathias Müller André Müller Shamsuddeen Hassan Muhammad Nanda Muhammad Ayanda Mnyakeni Jamshidbek Mirzakhalov Tapiwanashe Matangira Colin Leong Nze Lawson Sneha Kudugunta Yacine Jernite Affiliation: Bayero University Kano, University of South Florida, Hugging Face, Mathias Jenny Orhan Firat Bonaventure F. P. Dossou Sakhile Dlamini Nisansa de Silva Sakine Çabuk Ballı Stella Biderman Affiliation: Jacobs University Bremen, University of Moratuwa, EleutherAI, Alessia Battisti Ahmed Baruwa Ankur Bapna Pallavi Baljekar Israel Abebe Azime Affiliation: RobotsMali, Intel Labs, University of Zambia, Google, AIMS-AMMI, Ayodele Awokoya Duygu Ataman Orevaoghene Ahia Affiliation: Obafemi Awolowo University, University of Ibadan, Instadeep, Oghenefego Ahia Sweta Agrawal Mofetoluwa Adeyemi Affiliation: University of Maryland, Defence Space Administration Abuja,
Abstract
With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with five major public datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, mC4). Lower-resource corpora have systematic issues: At least 15 corpora have no usable text, and a significant fraction contains less than 50% sentences of acceptable quality. In addition, many are mislabeled or use nonstandard/ambiguous language codes. We demonstrate that these issues are easy to detect even for non-proficient speakers, and supplement the human audit with automatic analyses. Finally, we recommend techniques to evaluate and improve multilingual corpora and discuss potential risks that come with low-quality data releases.
原文 arXiv:2103.12028;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2103.12028v4