AI and the Everything in the Whole Wide World Benchmark
Inioluwa Deborah Raji Mozilla Foundation, UC Berkeley \AndEmily M. Bender Department of Linguistics University of Washington \AndAmandalynne Paullada Department of Linguistics University of Washington \AndEmily Denton Google Research \AndAlex Hanna Google Research
Abstract
There is a tendency across different subfields in AI to valorize a small collection of influential benchmarks. These benchmarks operate as stand-ins for a range of anointed common problems that are frequently framed as foundational milestones on the path towards flexible and generalizable AI systems. State-of-the-art performance on these benchmarks is widely understood as indicative of progress towards these long-term goals. In this position paper, we explore the limits of such benchmarks in order to reveal the construct validity issues in their framing as the functionally “general” broad measures of progress they are set up to be.
中文速览
大规模通用AI评测基准(benchmark)正被整个领域过度神化:ImageNet被当作"视觉理解"的终极标尺,GLUE被视为"语言理解"的通用量规,模型在这些榜单上刷新纪录就被默认等同于向通用智能迈进了一步。然而这些基准从根子上就存在"构念效度"(construct validity)问题——它们只是有限数据集与特定指标的组合,根本无法代表所声称的那种宽泛、普适的能力,就像一座宣称收藏了"世上万物"的博物馆,走到尽头打开最后那扇门,外面其实就是茫茫世界本身。这篇文章追溯了基准测试从上世纪80年代"公共任务框架"(Common Task Framework,CTF)的务实起源,到如今被滥用于评估模糊"通用能力"的演变轨迹,并以ImageNet和GLUE为案例,逐一拆解其在数据构成、指标设计和社区实践层面的局限。作者并非主张废弃这些基准,而是呼吁研究社区正视其评估范围的边界,避免用夸大的修辞掩盖实质性进展的缺失——因为继续将特定、有限、带有文化偏见的测试集供奉为通用能力的金标准,只会让整个领域在自我欺骗的道路上越走越远。
原文 arXiv:2111.15366;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2111.15366v1