KdConv: A Chinese Multi-domain Dialogue Dataset Towards Multi-turn Knowledge-driven Conversation
Hao Zhou Thanks: Equal contribution Chujie Zheng Kaili Huang Minlie Huang Thanks: Corresponding author: Minlie Huang. Xiaoyan ZhuConversational AI Group, AI Lab., Dept. of Computer Science, Tsinghua UniversityBeijing National Research Center for Information Science and Technology,
Abstract
The research of knowledge-driven conversational systems is largely limited due to the lack of dialog data which consist of multi-turn conversations on multiple topics and with knowledge annotations. In this paper, we propose a Chinese multi-domain knowledge-driven conversation dataset, KdConv, which grounds the topics in multi-turn conversations to knowledge graphs. Our corpus contains 4.5K conversations from three domains (film, music, and travel), and 86K utterances with an average turn number of 19.0. These conversations contain in-depth discussions on related topics and natural transition between multiple topics. To facilitate the following research on this corpus, we provide several benchmark models. Comparative results show that the models can be enhanced by introducing background knowledge, yet there is still a large space for leveraging knowledge to model multi-turn conversations for further research. Results also show that there are obvious performance differences between different domains, indicating that it is worth to further explore transfer learning and domain adaptation. The corpus and benchmark models are publicly available11 1 https://github.com/thu-coai/KdConv.
中文速览
人类式知识对话常受限于缺少既有多轮交流、又标注知识关联的多主题数据,KdConv因此构建了一个中文知识驱动对话数据集,覆盖电影、音乐和旅游三个领域。研究者先整理领域知识图谱,再由双方都能查阅知识的标注者进行无预设目标的多轮聊天,并为每句话标注相关知识事实,同时鼓励自然切换话题。数据集包含约4500段对话和8.6万条话语,平均每段19轮,话题可在1至4个之间深入转换;基准实验显示,引入背景知识能提升生成和检索模型的表现,但现有模型仍难以维持知识连贯性,且不同领域效果差异明显。它的重要性在于为多轮知识规划、知识衔接以及跨领域迁移研究提供了更贴近真实交流的中文数据和统一评测基础。
原文 arXiv:2004.04100;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2004.04100v1