CommanderSong: A Systematic Approach for Practical Adversarial Voice Recognition
Xuejing Yuan Affiliation: SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences, China Affiliation: School of Cyber Security, University of Chinese Academy of Sciences, China Yuxuan Chen Affiliation: Department of Computer Science, Florida Institute of Technology, USA Yue Zhao Affiliation: SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences, China Affiliation: School of Cyber Security, University of Chinese Academy of Sciences, China Yunhui Long Affiliation: Department of Computer Science, University of Illinois at Urbana-Champaign, USA Xiaokang Liu Affiliation: SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences, China Affiliation: School of Cyber Security, University of Chinese Academy of Sciences, China Kai Chen Thanks: Corresponding author: Affiliation: SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences, China Affiliation: School of Cyber Security, University of Chinese Academy of Sciences, China Shengzhi Zhang Affiliation: Department of Computer Science, Florida Institute of Technology, USA Affiliation: Department of Computer Science, Metropolitan College, Boston University, USA Heqing Huang XiaoFeng Wang Affiliation: School of Informatics and Computing, Indiana University Bloomington, USA Carl A. Gunter Affiliation: Department of Computer Science, University of Illinois at Urbana-Champaign, USA
Abstract
The popularity of automatic speech recognition (ASR) systems, like Google Assistant, Cortana, brings in security concerns, as demonstrated by recent attacks. The impacts of such threats, however, are less clear, since they are either less stealthy (producing noise-like voice commands) or requiring the physical presence of an attack device (using ultrasound speakers or transducers). In this paper, we demonstrate that not only are more practical and surreptitious attacks feasible but they can even be automatically constructed. Specifically, we find that the voice commands can be stealthily embedded into songs, which, when played, can effectively control the target system through ASR without being noticed. For this purpose, we developed novel techniques that address a key technical challenge: integrating the commands into a song in a way that can be effectively recognized by ASR through the air, in the presence of background noise, while not being detected by a human listener. Our research shows that this can be done automatically against real world ASR applications11 1 Demos of attacks are uploaded on the website (https://sites.google.com/view/commandersong/). We also demonstrate tha
原文 arXiv:1801.08535;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1801.08535v3