AI 中文总结
本研究针对端到端语音语言模型提出基于扰动的拒绝服务攻击,通过优化声学扰动抑制EOS生成以延长解码,在三类开源模型上验证了其攻击有效性及安全风险。
AI 中文摘要
多项研究表明,特制输入可诱导大语言模型(LLMs)生成过长输出,造成大量计算开销与资源消耗。现有多数拒绝服务(DoS)攻击仅针对纯文本LLMs,而端到端(E2E)语音LLMs正快速兴起;现有基于文本的DoS攻击主要依赖提示工程(如对抗后缀或语义诱导),利用文本输入的离散特性,无法直接迁移至连续语音输入。此外,此前语音模型安全研究主要聚焦于自动语音识别(ASR)或文本转语音(TTS)系统,端到端语音LLMs的DoS漏洞在很大程度上未被探索。为填补这一空白,我们提出针对E2E语音模型的基于扰动的DoS攻击:该方法不通过提示操纵诱导长输出,而是优化不可感知的声学扰动,在保留原始输入长度的同时直接影响模型的自回归生成过程。具体而言,我们将攻击构建为复合优化目标,通过整合加权的EOS损失、top-k logit损失、长度损失与语义对齐损失,联合抑制EOS生成、鼓励延长解码并大幅保留语义一致性;为进一步提升隐蔽性,我们采用语音活动检测(VAD)仅在浊音区域注入扰动。在三个开源E2E语音LLMs上开展的大量实验表明,我们的方法可实现稳定的攻击成功率,同时显著增加生成长度与GPU资源消耗,揭示了现代ALLMs的安全风险。
英文摘要
Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead and resource consumption. While most existing denial-of-service (DoS) attacks target text-only LLMs, end-to-end (E2E) speech LLMs are rapidly emerging. Existing text-based DoS attacks primarily rely on prompt engineering, such as adversarial suffixes or semantic inducement, which exploit the discrete nature of text inputs and therefore cannot be directly transferred to continuous speech inputs. Moreover, prior studies on speech model security mainly focus on ASR or TTS systems, leaving the DoS vulnerability of E2E speech LLMs largely unexplored. To address this gap, we propose the perturbation-based DoS attack targeting E2E speech models. Instead of inducing long outputs through prompt manipulation, our method optimizes imperceptible acoustic perturbations to directly influence the model's autoregressive generation process while preserving the original input length. Specifically, we formulate the attack as a composite optimization objective that jointly suppresses EOS generation, encourages prolonged decoding, and largely preserves semantic consistency by integrating weighted EOS loss, top-k logit loss, length loss, and semantic alignment loss. To further improve stealthiness, we employ voice activity detection (VAD) to inject perturbations only into voiced regions. Extensive experiments on three open-source E2E speech LLMs demonstrate that our method achieves stable attack success rate while significantly increasing generation length and GPU resource consumption, revealing security risks in modern ALLMs.