How would you say that? Pictures elicit better NLG data from the crowd.
Jekaterina Novikova, Oliver Lemon Verena Rieser Interaction Lab Heriot-Watt University Edinburgh, EH14 4AS, UK {j.novikova, o.lemon,
Abstract
Recent advances in corpus-based Natural Language Generation (NLG) hold the promise of being easily portable across domains, but require costly training data, consisting of meaning representations (MRs) paired with Natural Language (NL) utterances. In this work, we propose a novel framework for crowd-sourcing high quality NLG training data, using automatic quality control measures and evaluating different MRs with which to elicit data. We show that pictorial MRs result in better NL data being collected than logic-based MRs: utterances elicited by pictorial MRs are judged as significantly more natural, more informative, and better phrased, with a significant increase in average quality ratings (around 0.5 points on a 6-point scale), compared to using the logical MRs. As the MR becomes more complex, the benefits of pictorial stimuli increase. The collected data will be released as part of this submission.
中文速览
用图片代替逻辑符号来标注训练数据,能让众包工作者写出更自然的句子。自然语言生成(Natural Language Generation, NLG)系统需要大量"语义表示—自然语言"配对数据才能训练,但这类数据既昂贵又耗时。研究者提出了一套众包数据采集框架,核心创新是用图标拼成的图片(pictorial MR)替代传统的逻辑符号形式(如 inform(type[hotel],pricerange[expensive]))来引导工作者造句,并配合自动预校验和人工质量评分双重质控机制。实验共收集1410条语料,结果显示图片语义表示能显著提升数据质量——在6分量表上平均得分高出约0.5分,工作者写出的句子在自然度、信息量和表达方式三个维度上均明显更好,且语义越复杂优势越突出。这项工作为低成本、可跨领域扩展地构建NLG训练数据提供了切实可行的新路径。
原文 arXiv:1608.00339;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1608.00339v1