Question Answering is a Format; When is it Useful?
Matt Gardner♠♠\spadesuit, Jonathan Berant♠♠\spadesuit,♣♣\clubsuit, Hannaneh Hajishirzi♠♠\spadesuit,♢♢\diamondsuit, Alon Talmor♣♣\clubsuit, and Sewon Min♢♢\diamondsuit ♠♠\spadesuitAllen Institute for Artificial Intelligence ♣♣\clubsuitTel Aviv University ♢♢\diamondsuitUniversity of Washington
Abstract
Recent years have seen a dramatic expansion of tasks and datasets posed as question answering, from reading comprehension, semantic role labeling, and even machine translation, to image and video understanding. With this expansion, there are many differing views on the utility and definition of “question answering” itself. Some argue that its scope should be narrow, or broad, or that it is overused in datasets today. In this opinion piece, we argue that question answering should be considered a format which is sometimes useful for studying particular phenomena, not a phenomenon or task in itself. We discuss when a task is correctly described as question answering, and when a task is usefully posed as question answering, instead of using some other format.
中文速览
越来越多的NLP和视觉任务都在以"问答"(Question Answering, QA)形式呈现,导致这个词被滥用、含义模糊,研究者们既不清楚它究竟指什么,也不确定什么时候该用它。这篇文章提出,问答不是一种"任务"或"现象",而只是一种**格式**——就像分类、自然语言推理一样,是组织和呈现问题的一种方式。作者进一步给出了判断标准:只有当"理解问题本身的语言"是任务的实质组成部分时,使用问答格式才真正合适;反之,如果每个问题都能直接替换成一个整数编号而不改变任务本质,那本质上就是"槽填充"而非问答。基于这一框架,文章归纳了问答格式真正有价值的三类场景:满足人类真实信息需求、作为灵活的标注或探测机制、以及跨任务迁移模型参数,从而为社区提供了一套更清晰的语言,帮助大家判断何时该用、何时不该用问答格式,而不必对"又一个QA数据集"感到焦虑或一味追捧。
原文 arXiv:1909.11291;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1909.11291v1