A Span-Extraction Dataset for Chinese Machine Reading Comprehension
Yiming Cui Affiliation: Research Center for Social Computing and Information Retrieval (SCIR),Harbin Institute of Technology, Harbin, China Affiliation: State Key Laboratory of Cognitive Intelligence, iFLYTEK Research, China Affiliation: Ting Liu Affiliation: Research Center for Social Computing and Information Retrieval (SCIR),Harbin Institute of Technology, Harbin, China Affiliation: Wanxiang Che Affiliation: Research Center for Social Computing and Information Retrieval (SCIR),Harbin Institute of Technology, Harbin, China Affiliation: Li Xiao Affiliation: State Key Laboratory of Cognitive Intelligence, iFLYTEK Research, China Zhipeng Chen Affiliation: State Key Laboratory of Cognitive Intelligence, iFLYTEK Research, China Wentao Ma Affiliation: State Key Laboratory of Cognitive Intelligence, iFLYTEK Research, China Affiliation: iFLYTEK AI Research (Hebei), Langfang, China Affiliation: Shijin Wang Affiliation: State Key Laboratory of Cognitive Intelligence, iFLYTEK Research, China Affiliation: iFLYTEK AI Research (Hebei), Langfang, China Affiliation: Guoping Hu Affiliation: State Key Laboratory of Cognitive Intelligence, iFLYTEK Research, China
Abstract
Machine Reading Comprehension (MRC) has become enormously popular recently and has attracted a lot of attention. However, the existing reading comprehension datasets are mostly in English. In this paper, we introduce a Span-Extraction dataset for Chinese machine reading comprehension to add language diversities in this area. The dataset is composed by near 20,000 real questions annotated on Wikipedia paragraphs by human experts. We also annotated a challenge set which contains the questions that need comprehensive understanding and multi-sentence inference throughout the context. We present several baseline systems as well as anonymous submissions for demonstrating the difficulties in this dataset. With the release of the dataset, we hosted the Second Evaluation Workshop on Chinese Machine Reading Comprehension (CMRC 2018). We hope the release of the dataset could further accelerate the Chinese machine reading comprehension research.11 1 Resources are available: https://github.com/ymcui/cmrc2018.
原文 arXiv:1810.07366;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1810.07366v2