Measuring Attribution in Natural Language Generation Models
Hannah Rashkin,♣♢ Equal contribution. All authors contributed to all parts of the paper. ♠♠\spadesuit Led development of the conceptual framework. ♣♣\clubsuit Led human annotation study. ♢♢\diamondsuit Contributed to modeling experiments. ♡♡\heartsuit Provided project leadership and management. E-mail: Google Research Vitaly Nikolaev∗,♣♠ Google Research Matthew Lamm♠ Google Research Lora Aroyo♠ Google Research Michael Collins♠ Google Research Dipanjan Das♠♡ Google Research Slav Petrov♡ Google Research Gaurav Singh Tomar♢ Google Research Iulia Turc♢ Google Research David Reitter♠♡ Google Research
Abstract
With recent improvements in natural language generation (NLG) models for various applications, it has become imperative to have the means to identify and evaluate whether NLG output is only sharing verifiable information about the external world. In this work, we present a new evaluation framework entitled Attributable to Identified Sources (AIS) for assessing the output of natural language generation models, when such output pertains to the external world. We first define AIS and introduce a two-stage annotation pipeline for allowing annotators to appropriately evaluate model output according to AIS guidelines. We empirically validate this approach on generation datasets spanning three tasks (two conversational QA datasets, a summarization dataset, and a table-to-text dataset) via human evaluation studies that suggest that AIS could serve as a common framework for measuring whether model-generated statements are supported by underlying sources. We release guidelines for the human evaluation studies.
中文速览
自然语言生成(NLG)模型越来越容易"一本正经地胡说",却缺乏一套统一标准来衡量模型输出是否真正有据可查。为此,研究者提出了一个名为"可归因于已识别来源"(Attributable to Identified Sources,AIS)的评估框架,用一个直觉上简单的测试——"根据来源P,能否说出句子s"——来判断生成文本是否可以追溯到具体的信息来源,并借助"明示义"(explicature)这一语用学概念处理句子在上下文中的真实含义。研究者还设计了一套两阶段标注流程,先判断句子是否可解读,再判断其是否可归因,并在对话式问答、文本摘要和表格转文本三类任务上进行了人工评估实验,结果显示标注者之间能取得中到高度的一致性,不同模型的AIS得分也呈现出符合预期的差异。这项工作的意义在于,它为跨任务、可复现地评估NLG模型的"幻觉"和忠实性问题提供了一个统一、正式的框架,并公开了详细的标注指南,有望成为该领域的通用评测标准。
原文 arXiv:2112.12870;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2112.12870v2