Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing
Sarah Wiegreffe Thanks: Equal contributions. Affiliation: School of Interactive Computing Affiliation: Georgia Institute of Technology Email: Ana Marasović Affiliation: Allen Institute for AI Affiliation: University of Washington Email:
Abstract
Explainable Natural Language Processing (ExNLP) has increasingly focused on collecting human-annotated textual explanations. These explanations are used downstream in three ways: as data augmentation to improve performance on a predictive task, as supervision to train models to produce explanations for their predictions, and as a ground-truth to evaluate model-generated explanations. In this review, we identify 65 datasets with three predominant classes of textual explanations (highlights, free-text, and structured), organize the literature on annotating each type, identify strengths and shortcomings of existing collection methodologies, and give recommendations for collecting ExNLP datasets in the future.
原文 arXiv:2102.12060;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2102.12060v4