Soda: Million-scale Dialogue Distillation with Social Commonsense Contextualization
Hyunwoo Kim♡♠ Jack Hessel♡ Liwei Jiang♡♢ Peter West♢ Ximing Lu♢ Youngjae Yu♡ Pei Zhou♡♣ Ronan Le Bras♡ Malihe Alikhani† Gunhee Kim♠ Maarten Sap♡‡ Yejin Choi♡♢ ♡♡\heartsuit Allen Institute for Artificial Intelligence ♠♠\spadesuit Seoul National University ♢♢\diamondsuit University of Washington ♣♣\clubsuit University of Southern California ††\dagger University of Pittsburgh ‡‡\ddagger Carnegie Mellon University
Abstract
Data scarcity has been a long standing issue in the field of open-domain social dialogue. To quench this thirst, we present Soda: the first publicly available, million-scale high-quality social dialogue dataset. By contextualizing social commonsense knowledge from a knowledge graph, we are able to distill an exceptionally broad spectrum of social interactions from a large language model. Human evaluation shows that conversations in Soda are more consistent, specific, and (surprisingly) natural than those in prior human-authored datasets.
原文 arXiv:2212.10465;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2212.10465v3