The Dangers of Underclaiming: Reasons for Caution When Reporting How NLP Systems Fail
Samuel R. Bowman New York University
Abstract
Researchers in NLP often frame and discuss research results in ways that serve to deemphasize the field’s successes, often in response to the field’s widespread hype. Though well-meaning, this has yielded many misleading or false claims about the limits of our best technology. This is a problem, and it may be more serious than it looks: It harms our credibility in ways that can make it harder to mitigate present-day harms, like those involving biased systems for content moderation or resume screening. It also limits our ability to prepare for the potentially enormous impacts of more distant future advances. This paper urges researchers to be careful about these claims and suggests some research directions and communication strategies that will make it easier to avoid or rebut them.
中文速览
自然语言处理(NLP)领域长期存在"过度吹捧"的问题,但矫枉过正后,研究者开始走向另一个极端——系统性地低估模型的真实能力。作者将这种现象称为"低估"(underclaiming),具体表现为:用过时或较弱模型的失败结果来描述当前系统的局限、对理论证明做超出范围的延伸解读、以及将对抗性构造数据集上的低分当作真实能力的绝对度量。这些做法虽出于谨慎的好意,却导致学界对自身技术能力的共识长期落后于实际水平,既削弱了研究者在批评当下有害应用(如有偏见的内容审核、简历筛选系统)时的公信力,也妨碍了社会对更强大AI系统未来冲击的预判与准备。为此,作者呼吁研究者在写作和引用时更加严谨,并提出了一系列改进建议,包括引入评估新模型的规范、改进基准测试工具,以及发展模型性能预测的研究方向。
原文 arXiv:2110.08300;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2110.08300v3