Social Turing Tests: Crowdsourcing Sybil Detection
Gang Wang, Manish Mohanlal, Christo Wilson, Xiao Wang‡, Miriam Metzger†, Haitao Zheng and Ben Y. Zhao Department of Computer Science, U. C. Santa Barbara, CA USA †Department of Communications, U. C. Santa Barbara, CA USA ‡Renren Inc., Beijing, China
Abstract
As popular tools for spreading spam and malware, Sybils (or fake accounts) pose a serious threat to online communities such as Online Social Networks (OSNs). Today, sophisticated attackers are creating realistic Sybils that effectively befriend legitimate users, rendering most automated Sybil detection techniques ineffective. In this paper, we explore the feasibility of a crowdsourced Sybil detection system for OSNs. We conduct a large user study on the ability of humans to detect today’s Sybil accounts, using a large corpus of ground-truth Sybil accounts from the Facebook and Renren networks. We analyze detection accuracy by both “experts” and “turkers” under a variety of conditions, and find that while turkers vary significantly in their effectiveness, experts consistently produce near-optimal results. We use these results to drive the design of a multi-tier crowdsourcing Sybil detection system. Using our user study data, we show that this system is scalable, and can be highly effective either as a standalone system or as a complementary technique to current tools.
中文速览
伪造账号(Sybil账号)正在大规模入侵Facebook、人人网等社交平台,而现有自动检测算法因为伪造者越来越会伪装成真实用户而逐渐失效,研究者于是提出用"众包"的方式让真人来判断一个账号是否是假号。他们收集了来自Facebook(美国和印度)及人人网的大量已确认伪造账号,设计了一套用户实验,让"专家"、亚马逊众包平台工人(turker)和大学生分别对账号真伪做判断,系统地考察了准确率、人口统计特征、疲劳效应等因素的影响。结果发现,有经验的专家和大学生识别准确率接近最优且误判率极低,众包工人准确率稍低但整体仍表现良好,只是随着判断数量增加更容易出错。基于这些发现,他们设计了一套多层众包检测系统,模拟数据表明该系统在保持高准确率的同时具备规模可扩展性,成本也在可接受范围内——这为社交平台引入低成本、高效的人工智能协同伪号检测机制提供了重要的实证依据。
原文 arXiv:1205.3856;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1205.3856v2