Large datasets: A Pyrrhic win for computer vision?
Vinay Uday Prabhu UnifyID AI Labs Redwood City、Abeba Birhane School of Computer Science, UCD, Ireland Lero - The Irish Software Research Centre Equal contributions
Abstract
In this paper we investigate problematic practices and consequences of large scale vision datasets. We examine broad issues such as the question of consent and justice as well as specific concerns such as the inclusion of verifiably pornographic images in datasets. Taking the ImageNet-ILSVRC-2012 dataset as an example, we perform a cross-sectional model-based quantitative census covering factors such as age, gender, NSFW content scoring, class-wise accuracy, human-cardinality-analysis, and the semanticity of the image class information in order to statistically investigate the extent and subtleties of ethical transgressions. We then use the census to help hand-curate a look-up-table of images in the ImageNet-ILSVRC-2012 dataset that fall into the categories of verifiably pornographic: shot in a non-consensual setting (up-skirt), beach voyeuristic, and exposed private parts. We survey the landscape of harm and threats both society broadly and individuals face due to uncritical and ill-considered dataset curation practices. We then propose possible courses of correction and critique the pros and cons of these. We have duly open-sourced all of the code and the census meta-datasets gen
中文速览
大规模视觉数据集(如ImageNet)在收集数百万张人物图片时普遍缺乏当事人知情同意,由此引发了严重的伦理与隐私问题。研究者以ImageNet-ILSVRC-2012数据集为对象,系统开展了跨截面量化审计,涵盖年龄、性别、NSFW内容评分、分类准确率等多个维度,并手工整理出一份包含可核实色情内容、偷拍及暴露隐私部位图片的查找表。研究发现,数据集中不仅存在明确的违规色情图像,还借助反向图像搜索引擎可轻易还原受害者(多为女性)的真实身份,构成勒索、人肉搜索等现实威胁;与此同时,继承自WordNet的标签体系还将大量冒犯性词汇直接用于人物分类,加剧了对少数群体的伤害。这项工作的重要性在于,它用量化证据揭示了"为AI发展牺牲个体权益"这一长期被忽视的代价,并呼吁学界强制设立机构审查委员会(IRB)以规范大规模数据集的采集流程,所有代码与元数据集也已开源,供社区进一步研究使用。
原文 arXiv:2006.16923;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2006.16923v2