Automatic Detection of Fake News
Verónica Pérez-Rosas1, Bennett Kleinberg2, Alexandra Lefevre1 Rada Mihalcea1 1Computer Science and Engineering, University of Michigan 2Department of Psychology, University of Amsterdam
Abstract
The proliferation of misleading information in everyday access media outlets such as social media feeds, news blogs, and online newspapers have made it challenging to identify trustworthy news sources, thus increasing the need for computational tools able to provide insights into the reliability of online content. In this paper, we focus on the automatic identification of fake content in online news. Our contribution is twofold. First, we introduce two novel datasets for the task of fake news detection, covering seven different news domains. We describe the collection, annotation, and validation process in detail and present several exploratory analysis on the identification of linguistic differences in fake and legitimate news content. Second, we conduct a set of learning experiments to build accurate fake news detectors. In addition, we provide comparative analyses of the automatic and manual identification of fake news.
中文速览
网络上的假新闻越来越多,自动识别它们成为迫切需求,但现有研究要么依赖讽刺性内容(如"洋葱报"),要么局限于单一领域(如政治),缺乏覆盖多领域的高质量数据集。研究者构建了两个新数据集:一个通过亚马逊众包平台(Amazon Mechanical Turk)让工人为六大新闻领域(体育、商业、娱乐、政治、科技、教育)的真实新闻撰写假版本,另一个直接从网络收集名人八卦领域的真假新闻对;在此基础上,他们提取N元语法、标点、心理语言学等语言特征,训练假新闻分类器,最高达到78%的准确率,并与人工识别结果进行了对比分析。这项工作的价值在于:它提供了跨领域、无讽刺混淆因素的假新闻基准数据集,同时揭示了真假新闻在语言风格上的可量化差异,为构建更通用的自动核查工具奠定了基础。
原文 arXiv:1708.07104;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1708.07104v1