Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review
Yansong Gao, Bao Gia Doan, Zhi Zhang, Siqi Ma, Jiliang Zhang, Anmin Fu, Surya Nepal, and Hyoungshick Kim This version might be updated.Y. Gao is with the School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing, China and Data61, CSIRO, Sydney, Australia. e-mail: Doan is with the School of Computer Science, The University of Adelaide, Adelaide, Australia. e-mail: Zhang, S. Nepal are with Data61, CSIRO, Sydney, Australia. e-mail: {zhi.zhang; Ma is with School of Information Technology and Electrical Engineering, The University of Queensland, Brisbane, Australia. e-mail: Zhang is with the College of Computer Science and Electronic Engineering, Hunan University, Changsha, China. e-mail: Fu is with the School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing, China. e-mail: Kim is with Department of Computer Science and Engineering, College of Computing, Sungkyunkwan University, South Korea and Data61, CSIRO, Sydney, Australia. e-mail:
Abstract
Backdoor attacks insert hidden associations or triggers to the deep learning models to override correct inference such as classification and make the system perform maliciously according to the attacker-chosen target while behaving normally in the absence of the trigger. As a new and rapidly evolving realistic attack, it could result in dire consequences, especially considering that the backdoor attack surfaces are broad. In 2019, the U.S. Army Research Office started soliciting countermeasures and launching TrojAI project, the National Institute of Standards and Technology has initialized a corresponding online competition accordingly.
中文速览
深度学习模型存在一种隐蔽威胁——后门攻击(backdoor attack),攻击者可以在模型训练阶段植入隐藏"触发器",使模型在正常输入时表现正常,一旦出现特定触发信号就会按攻击者意图做出错误决策,例如把停止标志识别成限速标牌。针对这一领域缺乏系统梳理的现状,本文从攻击者能力和机器学习流水线各阶段出发,将后门攻击面归纳为六类——代码污染、外包、预训练模型、数据收集、协同学习和部署后攻击,并将防御手段划分为盲目后门移除、离线检测、在线检测和事后移除四大类,逐一分析各方法的优缺点。研究还揭示了后门攻击的"正面用途",包括保护模型知识产权、设置蜜罐捕捉对抗样本攻击以及验证数据删除请求。综合来看,现有防御手段明显落后于攻击技术,尚无任何单一防御能够抵御所有类型的后门攻击,这项系统性综述为该领域指明了亟待突破的研究方向。
原文 arXiv:2007.10760;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2007.10760v3