Recent Advances in Adversarial Training for Adversarial Robustness
Tao Bai1111Contact Author Jinqi Luo1 Jun Zhao1 Bihan Wen1 Qian Wang2 1Nanyang Technological University, Singapore 2Wuhan University, China {bait0002, luoj0021, junzhao,
Abstract
Adversarial training is one of the most effective approaches to defending deep learning models against adversarial examples. Unlike other defense strategies, adversarial training aims to enhance the robustness of models intrinsically. During the last few years, adversarial training has been studied and discussed from various aspects. A variety of improvements and developments of adversarial training are proposed, which were, however, neglected in existing surveys. For the first time in this survey, we systematically review the recent progress on adversarial training for adversarial robustness with a novel taxonomy. Then we discuss the generalization problems in adversarial training from three perspectives and highlight the challenges which are not fully tackled. Finally, we present potential future directions.
中文速览
深度神经网络容易被精心设计的微小扰动——即对抗样本(adversarial examples)——欺骗,如何让模型本质上具备抵抗这类攻击的能力,是领域内的核心难题。对抗训练(adversarial training)通过在每轮训练中将对抗样本纳入训练数据,以"最小化最坏情况损失"的方式直接增强模型鲁棒性,被公认为目前最有效的防御手段。这篇综述首次系统梳理了对抗训练领域近年来的研究进展,提出了一套新的分类体系,涵盖对抗正则化、课程式训练、集成对抗训练等多条技术路线,并从泛化能力角度深入分析了现有方法的不足与挑战。理解这一领域的全貌对于推动深度学习在安全敏感场景(如自动驾驶、医疗影像)中的可靠部署具有重要意义,该综述也为后续研究指明了若干值得攻关的方向。
原文 arXiv:2102.01356;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2102.01356v5