Backdoor Learning: A Survey
Yiming Li, Yong Jiang, Zhifeng Li, Shu-Tao Xia Manuscript received xxx, xxx; revised xxx, xxx.Yiming Li is with Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China (email: Jiang and Shu-Tao Xia are with Tsinghua Shenzhen International Graduate School, Tsinghua University, and also with Research Center of Artificial Intelligence, Peng Cheng Laboratory, Shenzhen, China (e-mail: Li is with Tencent Data Platform, Shenzhen, China (email:
Abstract
Backdoor attack intends to embed hidden backdoor into deep neural networks (DNNs), so that the attacked models perform well on benign samples, whereas their predictions will be maliciously changed if the hidden backdoor is activated by attacker-specified triggers. This threat could happen when the training process is not fully controlled, such as training on third-party datasets or adopting third-party models, which poses a new and realistic threat. Although backdoor learning is an emerging and rapidly growing research area, its systematic review, however, remains blank. In this paper, we present the first comprehensive survey of this realm. We summarize and categorize existing backdoor attacks and defenses based on their characteristics, and provide a unified framework for analyzing poisoning-based backdoor attacks. Besides, we also analyze the relation between backdoor attacks and relevant fields ( $i.e.,$ adversarial attacks and data poisoning), and summarize widely adopted benchmark datasets. Finally, we briefly outline certain future research directions relying upon reviewed works. A curated list of backdoor-related resources is also available at https://github.com/THUYimingLi
中文速览
深度神经网络在被第三方提供数据、平台或预训练模型时,攻击者可以悄悄在训练过程中植入"后门(backdoor)",使模型在正常输入上表现正常,却在含有特定触发器(trigger)的输入上产生攻击者预设的错误输出。这篇综述系统梳理了后门学习领域的攻击与防御方法,提出了一套统一的分析框架,并按照攻击与防御各自的特性进行分类整理,同时厘清了后门攻击与对抗样本、数据投毒等相关领域的关系。研究涵盖了从图像分类到语音识别的多种任务场景,总结了常用基准数据集,并指出了当前防御手段在理论保证与实际效果之间的根本矛盾。这是该领域首篇系统性综述,为研究者提供了清晰的问题全貌和方法导图,对推动更安全可靠的深度学习系统具有重要参考价值。
原文 arXiv:2007.08745;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2007.08745v5