Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text
Sebastian Gehrmann \AndElizabeth Clark Google Research New York, NY {gehrmann, eaclark, \AndThibault Sellam
Abstract
Evaluation practices in natural language generation (NLG) have many known flaws, but improved evaluation approaches are rarely widely adopted. This issue has become more urgent, since neural NLG models have improved to the point where they can often no longer be distinguished based on the surface-level features that older metrics rely on. This paper surveys the issues with human and automatic model evaluations and with commonly used datasets in NLG that have been pointed out over the past 20 years. We summarize, categorize, and discuss how researchers have been addressing these issues and what their findings mean for the current state of model evaluations. Building on those insights, we lay out a long-term vision for NLG evaluation and propose concrete steps for researchers to improve their evaluation processes. Finally, we analyze 66 NLG papers from recent NLP conferences in how well they already follow these suggestions and identify which areas require more drastic changes to the status quo.
中文速览
自然语言生成(NLG)领域长期存在评估体系失效的问题:自动指标(如BLEU、ROUGE)只衡量输出文本与参考文本的表面相似度,人工评估又因流程不规范而难以复现,加上常用数据集覆盖面窄、以英语为主,导致模型能力被系统性高估。这篇综述梳理了过去20年间研究者对NLG评估方法和数据集的批评与改进尝试,归纳出当前评估的核心症结:现有方法擅长从一堆模型里选出"相对最好"的那个,却无法真实刻画这个模型究竟好到什么程度、坏在哪里。为此,作者提出了一套可操作的改进路径,核心是发布以"暴露模型缺陷"为导向的评估报告,配合多种互补自动指标、严格人工评估和公开数据,以便未来研究者复验和迭代。通过对66篇近期顶会NLG论文的系统分析,作者发现现有工作平均只满足了建议标准的27%,指出数据集文档缺失、评估结论缺乏支撑等最急需改变的方向,对整个NLG研究社区的评估规范化具有重要参考价值。
原文 arXiv:2202.06935;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2202.06935v1