What Will it Take to Fix Benchmarking in Natural Language Understanding?
Samuel R. Bowman New York University E. Dahl Google Research, Brain Team
Abstract
Evaluation for many natural language understanding (NLU) tasks is broken: Unreliable and biased systems score so highly on standard benchmarks that there is little room for researchers who develop better systems to demonstrate their improvements. The recent trend to abandon IID benchmarks in favor of adversarially-constructed, out-of-distribution test sets ensures that current models will perform poorly, but ultimately only obscures the abilities that we want our benchmarks to measure. In this position paper, we lay out four criteria that we argue NLU benchmarks should meet. We argue most current benchmarks fail at these criteria, and that adversarial data collection does not meaningfully address the causes of these failures. Instead, restoring a healthy evaluation ecosystem will require significant progress in the design of benchmark datasets, the reliability with which they are annotated, their size, and the ways they handle social bias.
中文速览
当前主流自然语言理解(NLU)评测基准已陷入两难困境:模型得分已逼近甚至超越人类标注水平,却仍在简单测试上频繁出错,说明高分并不代表真正理解语言。为此,一些研究者转向"对抗性数据收集"——专门挑模型答错的例子来构建测试集,但作者认为这只是治标不治本,反而会激励研究者开发"犯不同错误"而非"真正更好"的模型。本文提出评测基准应满足四条标准:有效性(真正测到目标语言能力)、标注一致性(减少噪声和歧义)、充足的统计效力(数据量足以分辨系统优劣),以及对社会偏见的约束(不鼓励有害偏见的模型)。作者指出现有基准几乎无一满足这四条,并呼吁社区通过改进数据设计、引入领域专家与众包协作、扩大数据规模、附加偏见评测子集等方式,从根本上重建健康的评测生态,而非绕过问题另起炉灶。
原文 arXiv:2104.02145;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2104.02145v3