Distillation as a Defense to Adversarial Perturbations against Deep Neural Networks
Nicolas Papernot1, Patrick McDaniel1, Xi Wu4, Somesh Jha4, and Ananthram Swami3 Affiliation: 1Department of Computer Science and Engineering, Penn State University Affiliation: 4Computer Sciences Department, University of Wisconsin-Madison Affiliation: 3United States Army Research Laboratory, Adelphi, Maryland Affiliation:
Abstract
Deep learning algorithms have been shown to perform extremely well on many classical machine learning problems. However, recent studies have shown that deep learning, like other machine learning techniques, is vulnerable to adversarial samples: inputs crafted to force a deep neural network (DNN) to provide adversary-selected outputs. Such attacks can seriously undermine the security of the system supported by the DNN, sometimes with devastating consequences. For example, autonomous vehicles can be crashed, illicit or illegal content can bypass content filters, or biometric authentication systems can be manipulated to allow improper access. In this work, we introduce a defensive mechanism called defensive distillation to reduce the effectiveness of adversarial samples on DNNs. We analytically investigate the generalizability and robustness properties granted by the use of defensive distillation when training DNNs. We also empirically study the effectiveness of our defense mechanisms on two DNNs placed in adversarial settings. The study shows that defensive distillation can reduce effectiveness of sample creation from 95% to less than 0.5% on a studied DNN. Such dramatic gains can be
原文 arXiv:1511.04508;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1511.04508v2