Robust Physical-World Attacks on Deep Learning Visual Classification
Kevin Eykholt These authors contributed equally. University of Michigan, Ann Arbor Ivan Evtimov* University of Washington Earlence Fernandes University of Washington Bo Li University of California, Berkeley Amir Rahmati Samsung Research America and Stony Brook University Chaowei Xiao University of Michigan, Ann Arbor Atul Prakash University of Michigan, Ann Arbor Tadayoshi Kohno University of Washington Dawn Song University of California, Berkeley
Abstract
Recent studies show that the state-of-the-art deep neural networks (DNNs) are vulnerable to adversarial examples, resulting from small-magnitude perturbations added to the input. Given that that emerging physical systems are using DNNs in safety-critical situations, adversarial examples could mislead these systems and cause dangerous situations. Therefore, understanding adversarial examples in the physical world is an important step towards developing resilient learning algorithms. We propose a general attack algorithm, Robust Physical Perturbations (RP2), to generate robust visual adversarial perturbations under different physical conditions. Using the real-world case of road sign classification, we show that adversarial examples generated using RP2 achieve high targeted misclassification rates against standard-architecture road sign classifiers in the physical world under various environmental conditions, including viewpoints. Due to the current lack of a standardized testing method, we propose a two-stage evaluation methodology for robust physical adversarial examples consisting of lab and field tests. Using this methodology, we evaluate the efficacy of physical adversarial mani
中文速览
深度神经网络(DNN)在自动驾驶等安全关键场景中被广泛部署,但它们容易受到对抗样本(adversarial examples)的欺骗——只需对输入图像做微小扰动,就能让模型给出完全错误的判断。研究团队提出了一种名为"鲁棒物理扰动"(Robust Physical Perturbations,RP2)的攻击算法,通过在优化过程中模拟真实拍摄时的距离、角度等物理变化,生成能在现实环境中持续有效的视觉对抗扰动,并将扰动范围限制在物体本身而非背景。以道路交通标志为实验对象,他们将精心设计的黑白贴纸贴在真实停车标志上,在实验室静态测试中达到100%的误分类率,在行驶车辆拍摄的视频帧中也达到84.8%的攻击成功率,使分类器将停车标志误认为限速45标志。这项研究揭示了现实世界中物理对抗攻击对自动驾驶视觉系统的严峻威胁,为构建更鲁棒的深度学习模型提供了重要的安全警示。
原文 arXiv:1707.08945;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1707.08945v5