Soda: Million-scale Dialogue Distillation with Social Commonsense Contextualization
Hyunwoo Kim♡♠ Jack Hessel♡ Liwei Jiang♡♢ Peter West♢ Ximing Lu♢ Youngjae Yu♡ Pei Zhou♡♣ Ronan Le Bras♡ Malihe Alikhani† Gunhee Kim♠ Maarten Sap♡‡ Yejin Choi♡♢ ♡♡\heartsuit Allen Institute for Artificial Intelligence ♠♠\spadesuit Seoul National University ♢♢\diamondsuit University of Washington ♣♣\clubsuit University of Southern California ††\dagger University of Pittsburgh ‡‡\ddagger Carnegie Mellon University
Abstract
Data scarcity has been a long standing issue in the field of open-domain social dialogue. To quench this thirst, we present Soda: the first publicly available, million-scale high-quality social dialogue dataset. By contextualizing social commonsense knowledge from a knowledge graph, we are able to distill an exceptionally broad spectrum of social interactions from a large language model. Human evaluation shows that conversations in Soda are more consistent, specific, and (surprisingly) natural than those in prior human-authored datasets.
中文速览
日常社交对话数据极度稀缺,严重制约了对话系统的研发,为此研究团队提出了一套名为CO3的框架:先从常识知识图谱中抽取社会常识三元组,将其转化为叙事段落,再用GPT-3.5蒸馏生成自然对话,最终构建出包含150万段对话、超过1100万条语句的大规模社交对话数据集SODA。人工评测表明,SODA在自然度、一致性和具体性等维度上全面超越了此前由人工众包创作的对话数据集,打破了"机器生成必然逊于人工标注"的固有认知。基于SODA训练出的对话模型COSMO在从未见过的数据集上也能保持出色的泛化能力,与BlenderBot、Koala、Vicuna等主流模型的对比中平均胜率超过40%,部分场景下甚至优于人工撰写的参考回复。这项工作不仅为开放域对话研究提供了迄今最大的公开训练资源,还揭示了当前以知识问答为导向的大语言模型在闲聊自然度上的明显短板,为社交对话系统的后续研究提供了重要参照。
原文 arXiv:2212.10465;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2212.10465v3