More Than Reading Comprehension: A Survey on Datasets and Metrics of Textual Question Answering
Yang Bai, Daisy Zhe Wang Y. Bai and D.Z.Wang are with the Department of Computer and Information of Science and Engineering, University of Florida, Gainesville, FL, 32611. E-mail:
Abstract
Textual Question Answering (QA) aims to provide precise answers to users’ questions in natural language using unstructured data. One of the most popular approaches to this goal is machine reading comprehension(MRC). In recent years, many novel datasets and evaluation metrics based on classical MRC tasks have been proposed for broader textual QA tasks. In this paper, we survey 47 recent textual QA benchmark datasets and propose a new taxonomy from an application point of view. In addition, We summarize 8 evaluation metrics of textual QA tasks. Finally, we discuss current trends in constructing textual QA benchmarks and suggest directions for future work.
中文速览
文本问答(Textual QA)领域数据集繁多、任务形式各异,新入门的研究者往往难以厘清全貌,这篇论文对此做了一次系统性梳理。作者调研了47个近年来的文本问答基准数据集,从应用场景角度提出了一套新的分类体系,将其划分为经典机器阅读理解(MRC)、对话式问答、多跳推理问答、长答案问答、跨语言问答、开放域问答和常识问答共七大子任务,并系统总结了8种常用评测指标。调研结果显示,该领域正朝着更复杂推理、更多样语言和更开放答案形式的方向演进,当前模型在早期经典数据集上已接近甚至超越人类水平,但在多跳推理、常识理解等挑战性任务上仍有较大差距。这项工作为研究者提供了一份覆盖面广、对比维度丰富的参考指南,有助于快速了解文本问答领域的现状与未来研究方向。
原文 arXiv:2109.12264;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2109.12264v2