\DATASET: A Dataset of 1.3 Million Summaries with Diverse Extractive Strategies
Max Grusky1,2, Mor Naaman2, Yoav Artzi1,2 1Department of Computer Science, 2Cornell Tech Cornell University, New York, NY 10044
Abstract
We present \DATASET, a summarization dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications. Extracted from search and social media metadata between 1998 and 2017, these high-quality summaries demonstrate high diversity of summarization styles. In particular, the summaries combine abstractive and extractive strategies, borrowing words and phrases from articles at varying rates. We analyze the extraction strategies used in \DATASETsummaries against other datasets to quantify the diversity and difficulty of our new data, and train existing methods on the data to evaluate its utility and challenges.
中文速览
新闻摘要数据集长期面临数据规模小、风格单一的困境,现有数据集要么体量不足,要么用标题、子弹点等"凑数"手段替代真正的摘要。这项工作从38家主流新闻媒体抓取了近20年的HTML元数据,构建了一个包含130万篇文章及其配套摘要的数据集NEWSROOM,这些摘要均由记者和编辑亲笔撰写,原本用于搜索引擎和社交媒体展示。研究者设计了"抽取片段覆盖率(Coverage)"和"抽取片段密度(Density)"等指标,发现该数据集横跨从高度抽取式到高度生成式的完整摘要风格谱系,而现有主流数据集则明显偏向抽取式风格。在此数据集上对多个基线摘要模型进行评测后发现,无论是自动指标还是人工评估,现有方法都面临显著挑战,说明NEWSROOM既是推动数据密集型方法进步的重要资源,也为摘要研究提出了更真实、更全面的测试场景。
原文 arXiv:1804.11283;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1804.11283v2