Are we done with ImageNet?
Lucas Beyer1 Olivier J. Hénaff22\,{}^{2} Alexander Kolesnikov Xiaohua Zhai Aäron van den Oord 1Google Brain (Zürich, CH) and 2DeepMind (London, UK) All authors contributed equally, project led by first author. Correspondence: {lbeyer, henaff, akolesnikov, xzhai,
Abstract
Yes, and no. We ask whether recent progress on the ImageNet classification benchmark continues to represent meaningful generalization, or whether the community has started to overfit to the idiosyncrasies of its labeling procedure. We therefore develop a significantly more robust procedure for collecting human annotations of the ImageNet validation set. Using these new labels, we reassess the accuracy of recently proposed ImageNet classifiers, and find their gains to be substantially smaller than those reported on the original labels. Furthermore, we find the original ImageNet labels to no longer be the best predictors of this independently-collected set, indicating that their usefulness in evaluating vision models may be nearing an end. Nevertheless, we find our annotation procedure to have largely remedied the errors in the original labels, reinforcing ImageNet as a powerful benchmark for future research in visual recognition333The new labels and rater answers are available at https://github.com/google-research/reassessed-imagenet.
中文速览
ImageNet长期以来是衡量图像识别进展的核心基准,但近年来模型在上面的持续刷分,究竟代表真正的泛化能力提升,还是在"钻"这套标注流程的漏洞?研究者为此重新设计了一套更严格的人工标注方案,针对ImageNet验证集收集了新标签(ReaL Labels),解决了原标签"一图一标""备选标签过窄""类别边界模糊"等固有缺陷。用新标签重新评估后发现,早期模型在ImageNet上的进步能几乎同等幅度地体现在新标签上,但近期模型的实际进步幅度远比原指标显示的小——部分"提升"只是对原标注流程特性的过拟合;更值得注意的是,已有若干顶尖模型在新标签上的表现超过了原始ImageNet标签本身的"准确率",意味着后者作为评估标准的有效性正在接近极限。这项工作既警示了社区不能盲目追求ImageNet数字,也通过更高质量的标注修正了大量原始错误,为视觉识别研究提供了一个更可信的评估基础。
原文 arXiv:2006.07159;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2006.07159v1