e-CARE: a New Dataset for Exploring Explainable Causal Reasoning
Li Du, Xiao Ding, Kai Xiong, Ting Liu, and Bing Qin Research Center for Social Computing and Information Retrieval Harbin Institute of Technology, China {ldu, xding, kxiong, Corresponding author
Abstract
Understanding causality has vital importance for various Natural Language Processing (NLP) applications. Beyond the labeled instances, conceptual explanations of the causality can provide deep understanding of the causal facts to facilitate the causal reasoning process. However, such explanation information still remains absent in existing causal reasoning resources. In this paper, we fill this gap by presenting a human-annotated explainable CAusal REasoning dataset (e-CARE), which contains over 21K causal reasoning questions, together with natural language formed explanations of the causal questions. Experimental results show that generating valid explanations for causal facts still remains especially challenging for the state-of-the-art models, and the explanation information can be helpful for promoting the accuracy and stability of causal reasoning models.
中文速览
因果推理能力对自然语言处理至关重要,但现有数据集只告诉模型"A导致B",却从不解释"为什么A会导致B",导致模型只学会了表面规律、稳定性差。为此,研究团队构建了一个名为e-CARE的可解释因果推理数据集,包含超过2.1万道多项选择题,每道题都附有人工标注的自然语言概念解释(例如"铜是良导体,所以用铜接触火焰手指会马上感到烫"),并提出了配套的解释生成任务和专门衡量解释质量的CEQ评价指标。实验结果表明,当前最强的预训练语言模型在生成有效概念解释方面仍十分吃力,但将解释信息引入训练过程确实能提升模型的推理准确率和跨数据集稳定性。这项工作填补了因果推理数据集在"可解释性"上的空白,为开发真正理解因果机制而非仅凭统计捷径作答的模型指明了方向。
原文 arXiv:2205.05849;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.05849v1