Feature Purification: How Adversarial Training Performs Robust Deep Learning Thanks: V1 of this paper was presented at IAS on this date: https://video.ias.edu/csdm/2020/0316-YuanzhiLi. We polished writing and experiments in V1.5, V2 and V3. We added experiments showing that adversarial training can be done through low-rank updates in V4. We would like to thank Sanjeev Arora and Hadi Salman for many useful feedbacks and discussions. An extended abstract of this paper has appeared in FOCS 2021.
Zeyuan Allen-Zhu Email: Affiliation: Microsoft Research Redmond Yuanzhi Li Email: Affiliation: Carnegie Mellon University
Abstract
Despite the empirical success of using adversarial training to defend deep learning models against adversarial perturbations, so far, it still remains rather unclear what the principles are behind the existence of adversarial perturbations, and what adversarial training does to the neural network to remove them.
原文 arXiv:2005.10190;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2005.10190v4