DuplexJail:全双工模型在语音打断下安全对齐失效
DuplexJail: Spoken Interruption Attacks on Full-Duplex Speech Models
浏览论文内容
中文总结 AI 辅助
提出DuplexJail攻击,通过语音打断使全双工模型安全对齐失效,实验显示攻击成功率显著提升,揭示新的越狱向量。
中文摘要 AI 辅助
全双工语音模型在生成响应的同时接受用户语音,从而形成了一个尚未被充分探索的攻击面。我们提出了DuplexJail,它通过用户音频通道传递固定的、与请求无关的语音提示。我们比较了有害请求结束后固定延迟打断与模型流式文本中出现线索后触发拒答的打断两种策略。在来自AdvBench和HarmBench的四个开源模型和720个有害请求中,固定延迟打断使PersonaPlex在AdvBench上的整体响应攻击成功率提升至40.3%,PersonaPlex-RL提升至48.7%,分别增加了33.8和39.3个百分点。拒答触发策略分别达到35.6%和48.6%,所有试验均被评分,无论是否发生打断。选定条件也增加了FLM-Audio的有害响应率,而BayLing-Duplex则出现下降。这些发现将语音打断识别为一种越狱攻击向量,并促使在整个持续的全双工交互过程中评估安全性。
英文摘要
Full-duplex speech models accept user speech while generating responses, making input timing a potential safety concern. We introduce DuplexJail, which delivers fixed, request-independent spoken jailbreak prompts through the user audio channel. Across four open-source models and 720 harmful requests from AdvBench and HarmBench, we compare fixed-delay and refusal-triggered interruption with request-end and post-response controls. On AdvBench, Guided Completion at a 1.0 s delay raises whole-response attack success rates to 40.3% for PersonaPlex and 48.7% for PersonaPlex-RL, increases of 33.8 and 39.3 percentage points over baseline. On HarmBench, which was not used for prompt selection, the same prompt at a 0.5 s delay increases ASR by 14.7 and 12.3 points, respectively. Effects vary across models: selected conditions increase FLM-Audio's harmfulness, while BayLing-Duplex shows decreases. These results show that spoken-jailbreak effectiveness depends on delivery timing and motivate safety evaluation across stages of full-duplex interaction.
发表机构
- Dolby Laboratories(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。