MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling
Paweł Budzianowski1, Tsung-Hsien Wen2∗, Bo-Hsiang Tseng1, Iñigo Casanueva2∗, Stefan Ultes1, Osman Ramadan1 and Milica Gašić1 1Department of Engineering, University of Cambridge, UK, 2PolyAI, London, UK
Abstract
Even though machine learning has become the major scene in dialogue research community, the real breakthrough has been blocked by the scale of data available. To address this fundamental obstacle, we introduce the Multi-Domain Wizard-of-Oz dataset (MultiWOZ), a fully-labeled collection of human-human written conversations spanning over multiple domains and topics. At a size of $10$ k dialogues, it is at least one order of magnitude larger than all previous annotated task-oriented corpora. The contribution of this work apart from the open-sourced dataset labelled with dialogue belief states and dialogue actions is two-fold: firstly, a detailed description of the data collection procedure along with a summary of data structure and analysis is provided. The proposed data-collection pipeline is entirely based on crowd-sourcing without the need of hiring professional annotators; secondly, a set of benchmark results of belief tracking, dialogue act and response generation is reported, which shows the usability of the data and sets a baseline for future studies.
中文速览
任务导向对话系统长期受限于标注数据规模太小,难以支撑大规模机器学习研究。为此,研究者构建了 MultiWOZ 数据集——一个涵盖多个领域(如餐厅、酒店、景点、出租车、火车等)的大规模人人对话语料库,共约 1 万段对话,附带对话信念状态(dialogue belief state)和对话动作(dialogue act)的完整标注,规模比此前所有同类有标注任务导向语料库至少大一个数量级。数据采集基于"奥兹巫师"(Wizard-of-Oz)框架,通过众包平台招募普通网络工作者扮演用户和系统角色进行对话,无需专业标注人员,大幅降低了成本并保证了语言的自然性。在此数据集上,研究者还报告了信念追踪、对话动作预测和端到端回复生成等基准实验结果,验证了数据集的实用价值。MultiWOZ 的发布为多领域对话系统的数据驱动研究提供了一个量级上的飞跃,对推动端到端对话建模和模块化系统评测都具有重要意义。
原文 arXiv:1810.00278;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1810.00278v3