Datasheets for Datasets
Timnit Gebru Black in AI , Jamie Morgenstern University of Washington , Briana Vecchione Cornell University , Jennifer Wortman Vaughan Microsoft Research , Hanna Wallach Microsoft Research , Hal Daumé III Microsoft Research;University of Maryland and Kate Crawford Microsoft Research
Abstract
Data plays a critical role in machine learning. Every machine learning model is trained and evaluated using data, quite often in the form of static datasets. The characteristics of these datasets fundamentally influence a model’s behavior: a model is unlikely to perform well in the wild if its deployment context does not match its training or evaluation datasets, or if these datasets reflect unwanted societal biases. Mismatches like this can have especially severe consequences when machine learning models are used in high-stakes domains, such as criminal justice (Garvie et al., 2016; Systems, 2017; Andrews et al., 2006), hiring (Mann and O’Neil, 2016), critical infrastructure (O’Connor, 2017; Chui, 2017), and finance (Lin, 2012). Even in other domains, mismatches may lead to loss of revenue or public relations setbacks. Of particular concern are recent examples showing that machine learning models can reproduce or amplify unwanted societal biases reflected in training datasets (Buolamwini and Gebru, 2018; Dastin, 2018; Bolukbasi et al., 2016). For these and other reasons, the World Economic Forum suggests that all entities should document the provenance, creation, and use of machin
中文速览
机器学习模型的好坏高度依赖训练和评估数据集的质量,但长期以来业界缺乏记录数据集来源、构成与用途的统一标准,导致模型偏差、结果难以复现、数据被不当使用等问题屡见不鲜。受电子元器件随附规格说明书的启发,作者提出"数据集数据表"(datasheets for datasets)这一框架——要求每个数据集都附上一份结构化文档,涵盖创建动机、数据构成、采集过程、预处理方式、推荐用途、分发方式及维护计划等七大环节的具体问题。这套问题历经约两年打磨,经过多家科技公司产品团队的实际试用、法律团队审查以及数十位研究者与政策制定者的反馈后持续迭代完善,并以真实数据集为例展示了如何填写。该提案的意义在于:它既督促数据集创建者在整个数据生命周期中主动反思潜在风险与社会偏见,也让数据集使用者获得足够的信息来做出明智选择,从而推动机器学习领域在数据层面实现更高的透明度与问责性。
原文 arXiv:1803.09010;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1803.09010v8