Bringing the People Back In: Contesting Benchmark Machine Learning Datasets
Emily Denton Alex Hanna Razvan Amironesei Andrew Smart Hilary Nicole Morgan Klaus Scheuerman
Abstract
In response to algorithmic unfairness embedded in sociotechnical systems, significant attention has been focused on the contents of machine learning datasets which have revealed biases towards white, cisgender, male, and Western data subjects. In contrast, comparatively less attention has been paid to the histories, values, and norms embedded in such datasets. In this work, we outline a research program – a genealogy of machine learning data – for investigating how and why these datasets have been created, what and whose values influence the choices of data to collect, the contextual and contingent conditions of their creation. We describe the ways in which benchmark datasets in machine learning operate as infrastructure and pose four research questions for these datasets. This interrogation forces us to “bring the people back in” by aiding us in understanding the labor embedded in dataset construction, and thereby presenting new avenues of contestation for other researchers encountering the data.
中文速览
机器学习系统之所以对特定群体造成伤害,根源不只在于数据集里谁被代表、谁被忽视,更在于数据集是怎么被创建出来的、背后承载了哪些历史、价值观和权力关系——而这一层面长期被忽视。这篇文章借用福柯的"谱系学"(genealogy)方法,提出了一套研究纲领:把机器学习基准数据集(benchmark dataset)当作基础设施(infrastructure)来审视,围绕四个核心问题追问它们的创建动机、历史脉络、权威化过程以及内嵌的劳动与伦理关系。研究者主张把"人"重新带回对数据的分析中心,揭示那些在日常科研中被自然化、隐形化的主观决策和权力痕迹。这项工作的意义在于:它不止于追求透明度,更致力于为研究者提供真正可操作的"可质疑性"(contestability),让人们能够识别并挑战数据基础设施中被默认接受的假设,从而为更公正的AI系统建设打开新的路径。
原文 arXiv:2007.07399;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2007.07399v1