Automatic Detection of Fake News
Verónica Pérez-Rosas1, Bennett Kleinberg2, Alexandra Lefevre1 Rada Mihalcea1 1Computer Science and Engineering, University of Michigan 2Department of Psychology, University of Amsterdam
Abstract
The proliferation of misleading information in everyday access media outlets such as social media feeds, news blogs, and online newspapers have made it challenging to identify trustworthy news sources, thus increasing the need for computational tools able to provide insights into the reliability of online content. In this paper, we focus on the automatic identification of fake content in online news. Our contribution is twofold. First, we introduce two novel datasets for the task of fake news detection, covering seven different news domains. We describe the collection, annotation, and validation process in detail and present several exploratory analysis on the identification of linguistic differences in fake and legitimate news content. Second, we conduct a set of learning experiments to build accurate fake news detectors. In addition, we provide comparative analyses of the automatic and manual identification of fake news.
中文速览
社交媒体、新闻博客和在线报纸中假新闻泛滥,使人们难以判断网上内容是否可信,尤其需要能跨领域识别虚假新闻的工具。研究者构建了两个新数据集:一个通过众包把六类真实新闻改写成假新闻,另一个从网络收集名人领域的真假新闻,覆盖七个领域,并结合词语、标点和心理语言学等写作特征训练分类模型。实验发现,假新闻在语言表达上确实存在可识别差异,自动检测准确率最高达到78%,同时还与人工判断能力进行了比较。其重要性在于,它避开了讽刺新闻带来的幽默干扰和事实核查网站领域单一的问题,为更广泛、可复用的假新闻研究提供了数据和方法基础。
原文 arXiv:1708.07104;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1708.07104v1