SpeechGuard:针对语音识别模型后门攻击的在线防御
SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models
浏览论文内容
中文总结 AI 辅助
研究针对语音识别模型运行时的后门攻击,提出SpeechGuard在线防御管道,改进STRIP方法为S-STRIP检测过滤中毒样本,利用时频掩蔽等净化样本,经实验验证其能有效减轻后门威胁并维持预测准确率。
中文摘要 AI 辅助
后门攻击对神经网络模型构成严重威胁,攻击者可在训练阶段通过操控一小部分训练数据植入后门。在诸如自动驾驶语音交互等安全敏感应用中,后门攻击带来巨大安全风险。本研究针对语音识别模型运行时实施后门防御措施,考虑音频信号特征。我们提出SpeechGuard,首个旨在识别和净化中毒音频样本的在线后门防御管道。具体而言,改进STRIP方法进行自适应扰动注入以检测和过滤中毒样本,即S-STRIP。更重要的是,进一步考虑中毒样本净化。利用时频掩蔽抑制触发信号表达并基于自编码器自主生成掩码。两阶段处理防止模型中的后门被触发,即便输入携带触发信号的语音也能准确预测。大量实验表明SpeechGuard能准确过滤中毒样本,通过净化可显著减轻后门威胁并保持一定预测准确率。
英文摘要
Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as voice interaction for autonomous driving, the presence of backdoor attacks introduces substantial security risks. This study focuses on implementing backdoor defense measures for speech recognition models in run-time, taking into account the characteristics of audio signals. We propose SpeechGuard, the first online backdoor defense pipeline designed to identify and purify poisoned audio samples. Specifically, we improve STRIP method to perform adaptive perturbation injection to detect and filter poisoned samples, named as S-STRIP. More importantly, we further consider the purification of poisoned samples. We utilize time-frequency (T-F) masking to suppress the expression of trigger signals and autonomously generate masks based on an autoencoder. The two-stage processing prevents the backdoor in the model from being triggered, and even input speech carrying triggers can be accurately predicted. Extensive experimental demonstrate that SpeechGuard can accurately filter out poisoned samples. Through purification, it can significantly mitigate the backdoor threat while maintaining a certain prediction accuracy.