CommanderSong: A Systematic Approach for Practical Adversarial Voice Recognition
Xuejing Yuan SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences, China School of Cyber Security, University of Chinese Academy of Sciences, China Yuxuan Chen Department of Computer Science, Florida Institute of Technology, USA Yue Zhao SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences, China School of Cyber Security, University of Chinese Academy of Sciences, China Yunhui Long Department of Computer Science, University of Illinois at Urbana-Champaign, USA Xiaokang Liu SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences, China School of Cyber Security, University of Chinese Academy of Sciences, China Kai Chen Corresponding author: SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences, China School of Cyber Security, University of Chinese Academy of Sciences, China Shengzhi Zhang Department of Computer Science, Florida Institute of Technology, USA Department of Computer Science, Metropolitan College, Boston University, USA Heqing Huang XiaoFeng Wang School of Informatics and Computing, Indiana University Bloomington, USA Carl A. Gunter Department of Computer Science, University of Illinois at Urbana-Champaign, USA
Abstract
The popularity of automatic speech recognition (ASR) systems, like Google Assistant, Cortana, brings in security concerns, as demonstrated by recent attacks. The impacts of such threats, however, are less clear, since they are either less stealthy (producing noise-like voice commands) or requiring the physical presence of an attack device (using ultrasound speakers or transducers). In this paper, we demonstrate that not only are more practical and surreptitious attacks feasible but they can even be automatically constructed. Specifically, we find that the voice commands can be stealthily embedded into songs, which, when played, can effectively control the target system through ASR without being noticed. For this purpose, we developed novel techniques that address a key technical challenge: integrating the commands into a song in a way that can be effectively recognized by ASR through the air, in the presence of background noise, while not being detected by a human listener. Our research shows that this can be done automatically against real world ASR applications111Demos of attacks are uploaded on the website (https://sites.google.com/view/commandersong/). We also demonstrate that
中文速览
智能语音控制系统(如 Google Assistant、Siri)正被越来越多的人使用,但研究者发现它们面临一种此前从未被系统证明过的隐蔽威胁:攻击者能把恶意语音指令悄无声息地藏进一首普通歌曲里,让听歌的人毫无察觉,而麦克风背后的语音识别系统(ASR)却会将其解读为真实命令并执行。为此,研究团队开发了一套基于梯度下降的对抗样本生成方法,利用开源语音识别框架 Kaldi 的声学模型,在保持歌曲听感自然的前提下,将目标指令的声学特征融入歌曲音频,同时引入扬声器电子噪声模型以确保攻击在真实物理环境中依然有效。实验生成了逾 200 首"指令歌曲"(CommanderSong),对 Kaldi 的攻击成功率高达 96–100%,对商业系统讯飞(iFLYTEK)的黑盒迁移攻击同样奏效,且在 200 余名亚马逊众包用户参与的用户研究中无一人发现歌曲中隐藏的指令;这些歌曲还可经由 YouTube、广播传播,理论上可同时影响数以百万计的用户。这项研究表明针对主流 DNN 语音识别系统的实用隐蔽攻击已完全可行,提醒业界必须认真评估语音交互场景下的对抗安全风险。
原文 arXiv:1801.08535;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1801.08535v3