\ourdata: Information-Seeking Conversations with Mixed-Initiative Interactions
Zeqiu Wu ♠♠{}^{\spadesuit}start_FLOATSUPERSCRIPT ♠ end_FLOATSUPERSCRIPT Ryu Parish ♠♠{}^{\spadesuit}start_FLOATSUPERSCRIPT ♠ end_FLOATSUPERSCRIPT Hao Cheng ♣♣{}^{\clubsuit}start_FLOATSUPERSCRIPT ♣ end_FLOATSUPERSCRIPT Sewon Min ♠♠{}^{\spadesuit}start_FLOATSUPERSCRIPT ♠ end_FLOATSUPERSCRIPT Prithviraj Ammanabrolu ♢♢{}^{\diamondsuit}start_FLOATSUPERSCRIPT ♢ end_FLOATSUPERSCRIPT Mari Ostendorf ♠♠{}^{\spadesuit}start_FLOATSUPERSCRIPT ♠ end_FLOATSUPERSCRIPT Hannaneh Hajishirzi ♠♠{}^{\spadesuit}start_FLOATSUPERSCRIPT ♠ end_FLOATSUPERSCRIPT♢♢{}^{\diamondsuit}start_FLOATSUPERSCRIPT ♢ end_FLOATSUPERSCRIPT ♠♠{}^{\spadesuit}start_FLOATSUPERSCRIPT ♠ end_FLOATSUPERSCRIPTUniversity of Washington ♣♣{}^{\clubsuit}start_FLOATSUPERSCRIPT ♣ end_FLOATSUPERSCRIPTMicrosoft Research ♢♢{}^{\diamondsuit}start_FLOATSUPERSCRIPT ♢ end_FLOATSUPERSCRIPTAllen Institute for AI
Abstract
In an information-seeking conversation, a user may ask questions that are under-specified or unanswerable. An ideal agent would interact by initiating different response types according to the available knowledge sources. However, most current studies either fail to or artificially incorporate such agent-side initiative. This work presents \ourdata, a dataset for Information-Seeking Conversations with mixed-initiative Interactions. It contains 4.7K user-agent turns from 805 human-human conversations where the agent searches over Wikipedia and either directly answers, asks for clarification, or provides relevant information to address user queries. The data supports two subtasks, evidence passage identification and response generation, as well as a human evaluation protocol to assess model performance. We report results of two systems based on state-of-the-art models of conversational knowledge identification and open-domain question answering. Both systems significantly underperform humans, suggesting ample room for improvement in future studies.111We open-source all data and code at https://github.com/ellenmellon/INSCIT.
中文速览
真实的信息查询对话中,用户提问常常模糊不清或无法直接回答,理想的对话系统应能灵活地直接作答、反问澄清或提供相关部分信息,而非千篇一律地"给答案"或"无答案"——现有数据集和模型在这方面普遍缺失。为此,研究团队构建了INSCIT数据集,通过众包收集了805段真实的人人对话、共4700余轮,对话中的"智能体"一侧在搜索维基百科后会根据实际情况选择直接回答(72%)、提出澄清问题(13%)或给出相关但不完整的信息(13%),并为每条回答标注了支撑证据段落。基于此数据集,团队设计了证据段落识别和回复生成两项子任务,并给出了基于当前最优开放域问答模型的两条基线,结果显示两者均大幅落后于人类水平,尤其在需要澄清或提供相关信息的复杂场景下差距最为明显。这项工作为推动更自然、更具主动性的对话式信息检索系统研究提供了重要的数据基础和评测框架。
原文 arXiv:2207.00746;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2207.00746v2