The E2E NLG Shared Task
Jekaterina Novikova, Ondřej Dušek and Verena Rieser School of Mathematical and Computer Sciences Heriot-Watt University, Edinburgh j.novikova, o.dusek,
Abstract
This paper describes the E2E data, a new dataset for training end-to-end, data-driven natural language generation systems in the restaurant domain, which is ten times bigger than existing, frequently used datasets in this area. The E2E dataset poses new challenges: (1) its human reference texts show more lexical richness and syntactic variation, including discourse phenomena; (2) generating from this set requires content selection. As such, learning from this dataset promises more natural, varied and less template-like system utterances. We also establish a baseline on this dataset, which illustrates some of the difficulties associated with this data.
中文速览
端到端自然语言生成(end-to-end NLG)领域长期依赖规模小、词汇固定的数据集,导致系统输出千篇一律、缺乏自然语言应有的丰富性。为此,研究者利用众包方式收集了餐厅领域的E2E数据集,包含超过5万条意义表示与自然语言参考文本的配对,规模是同类常用数据集的十倍,且采用图片而非文字提示来激发标注者写出更自然、更丰富的描述。分析表明,该数据集在词汇多样性、句法复杂度和篇章现象上均显著优于已有数据集,还引入了内容选择这一额外挑战——生成系统需自行判断哪些属性值值得表达。以TGen序列到序列模型为基线的实验表明,更大的数据规模和更多的参考文本确实有助于提升自动评估指标,但内容选择和开放词汇等问题仍未解决。这一数据集的公开发布为训练能够产生更自然、更多样输出的生成系统提供了重要资源,同时也为领域内树立了更高的研究挑战标准。
原文 arXiv:1706.09254;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1706.09254v2