A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis
Shai Gretz, Roni Friedman∗, Edo Cohen-Karlik∗, Assaf Toledo∗, Dan Lahav, Ranit Aharonov and Noam Slonim IBM Research These authors equally contributed to this work. Shai Gretz, Roni Friedman∗, Edo Cohen-Karlik∗, Assaf Toledo∗, Dan Lahav, Ranit Aharonov and Noam Slonim IBM Research These authors equally contributed to this work.
Abstract
Identifying the quality of free-text arguments has become an important task in the rapidly expanding field of computational argumentation. In this work, we explore the challenging task of argument quality ranking. To this end, we created a corpus of $30{,}497$ arguments carefully annotated for point-wise quality, released as part of this work. To the best of our knowledge, this is the largest dataset annotated for point-wise argument quality, larger by a factor of five than previously released datasets. Moreover, we address the core issue of inducing a labeled score from crowd annotations by performing a comprehensive evaluation of different approaches to this problem. In addition, we analyze the quality dimensions that characterize this dataset. Finally, we present a neural method for argument quality ranking, which outperforms several baselines on our own dataset, as well as previous methods published for another dataset.
中文速览
围绕"如何自动评估自由文本论点(argument)的质量"这一难题,研究者构建了迄今最大的论点质量标注数据集 IBM-Rank-30k,包含约 3 万条论点,比此前同类数据集大五倍。为了从众多标注者的二元投票中提炼出连续质量分数,论文系统比较了多种评分方法,发现结合标注者可信度的加权平均(Weighted-Average)与 MACE 概率模型各有优劣,并深入分析了不同评分方案对模型训练和评估的实际影响。在此基础上,论文提出了一种基于 BERT 的神经排名模型,在自建数据集和已有公开数据集上均超越了多个基线方法及已有最优方法。这项工作既为计算论辩领域提供了宝贵的大规模标注资源,也为自动决策支持、论点检索和写作辅助等实际应用奠定了更坚实的技术基础。
原文 arXiv:1911.11408;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1911.11408v1