Attribute-Guided Adversarial Training for Robustness to Natural Perturbations
Tejas Gokhale Note: Work performed during internship at LLNL. Rushil Anirudh Bhavya Kailkhura Jayaraman J. Thiagarajan Chitta Baral Yezhou Yang
Abstract
While existing work in robust deep learning has focused on small pixel-level norm-based perturbations, this may not account for perturbations encountered in several real-world settings. In many such cases although test data might not be available, broad specifications about the types of perturbations (such as an unknown degree of rotation) may be known. We consider a setup where robustness is expected over an unseen test domain that is not i.i.d. but deviates from the training domain. While this deviation may not be exactly known, its broad characterization is specified a priori, in terms of attributes. We propose an adversarial training approach which learns to generate new samples so as to maximize exposure of the classifier to the attributes-space, without having access to the data from the test domain. Our adversarial training solves a min-max optimization problem, with the inner maximization generating adversarial perturbations, and the outer minimization finding model parameters by optimizing the loss on adversarial perturbations generated from the inner maximization. We demonstrate the applicability of our approach on three types of naturally occurring perturbations — object
原文 arXiv:2012.01806;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2012.01806v3