Maven-Ere: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction
Xiaozhi Wang Thanks: indicates equal contribution. Affiliation: Department of Computer Science and Technology, BNRist; Yulin Chen Affiliation: Shenzhen International Graduate School; Ning Ding Affiliation: Department of Computer Science and Technology, BNRist; Hao Peng Affiliation: Department of Computer Science and Technology, BNRist; Affiliation: Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China Zimu Wang Affiliation: Xi’an Jiaotong-Liverpool University, Suzhou, China Yankai Lin Thanks: Partly done while Y.Lin and P.Li were at Tencent. Affiliation: Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China Affiliation: Beijing Key Laboratory of Big Data Management and Analysis Methods, Beijing, China Xu Han, Lei Hou, Juanzi Li , Zhiyuan Liu Thanks: Corresponding author: J.Li. Affiliation: Department of Computer Science and Technology, BNRist; Affiliation: Department of Computer Science and Technology, BNRist; Affiliation: Department of Computer Science and Technology, BNRist; Affiliation: Department of Computer Science and Technology, BNRist; Affiliation: THU-Siemens Ltd., China Joint Research Center for Industrial Intelligence and IoT; Peng Li, Jie Zhou Affiliation: Pattern Recognition Center, WeChat AI, Tencent Inc,
Abstract
The diverse relationships among real-world events, including coreference, temporal, causal, and subevent relations, are fundamental to understanding natural languages. However, two drawbacks of existing datasets limit event relation extraction (ERE) tasks: (1) Small scale. Due to the annotation complexity, the data scale of existing datasets is limited, which cannot well train and evaluate data-hungry models. (2) Absence of unified annotation. Different types of event relations naturally interact with each other, but existing datasets only cover limited relation types at once, which prevents models from taking full advantage of relation interactions. To address these issues, we construct a unified large-scale human-annotated ERE dataset Maven-Ere with improved annotation schemes. It contains $103,193$ event coreference chains, $1,216,217$ temporal relations, $57,992$ causal relations, and $15,841$ subevent relations, which is larger than existing datasets of all the ERE tasks by at least an order of magnitude. Experiments show that ERE on Maven-Ere is quite challenging, and considering relation interactions with joint learning can improve performances. The dataset and source codes
原文 arXiv:2211.07342;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2211.07342v1