The Benchmark Lottery
Mostafa Dehghani Google Brain \AndYi Tay∗ Google Research \AndAlexey A. Gritsenko∗ Google Brain \AndZhe Zhao Google Brain \AndNeil Houlsby Google Brain \AndFernando Diaz Google Brain \AndDonald Metzler† Google Research \AndOriol Vinyals† DeepMind Equal contribution, †equal advising.
Abstract
The world of empirical machine learning (ML) strongly relies on benchmarks in order to determine the relative effectiveness of different algorithms and methods. This paper proposes the notion of a benchmark lottery that describes the overall fragility of the ML benchmarking process. The benchmark lottery postulates that many factors, other than fundamental algorithmic superiority, may lead to a method being perceived as superior. On multiple benchmark setups that are prevalent in the ML community, we show that the relative performance of algorithms may be altered significantly simply by choosing different benchmark tasks, highlighting the fragility of the current paradigms and potential fallacious interpretation derived from benchmarking ML methods. Given that every benchmark makes a statement about what it perceives to be important, we argue that this might lead to biased progress in the community. We discuss the implications of the observed phenomena and provide recommendations on mitigating them using multiple machine learning domains and communities as use cases, including natural language processing, computer vision, information retrieval, recommender systems, and reinforcemen
中文速览
机器学习领域的进步高度依赖基准测试(benchmark),但基准本身存在严重的脆弱性——一个方法能否被视为"更优",往往并非因为它在算法上真的更强,而是因为它恰好与社区选定的那套任务高度契合。作者将这一现象称为"基准彩票"(benchmark lottery),并通过自然语言处理、计算机视觉、信息检索、推荐系统和强化学习等多个领域的实证分析,系统地展示了任务选择偏差(task selection bias)如何显著改变不同算法的相对排名。研究还指出,基准并非静态的,它随时间演化,某些领域缺乏统一标准甚至让研究者得以"反向操作"——量身定制实验来迎合自己的模型,从而变相"操控"彩票结果。这项工作呼吁整个机器学习社区对基准的选取、聚合和复用方式保持更清醒的认识,并提出了若干改进建议,对于避免社区资源和研究方向被基准体系无意识地引向偏路具有重要的警示价值。
原文 arXiv:2107.07002;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2107.07002v1