arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

音频大语言模型中基于韵律的越狱:一项对照研究与机制分析

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

Jiachen Qian, Junyu Li

arXiv 2607.26541首次发表:更新:

发表机构

City University of Hong Kong; City University of Hong Kong (Dongguan)(香港城市大学; 香港城市大学(东莞))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过控制文本内容改变语音传递预设构建评估协议与基准,发现特定韵律特征的语音能提升音频大语言模型的越狱成功率,表明语音传递是音频大语言模型安全评估的重要因素。

AI 中文摘要

具备音频能力的基础模型支持端到端口语交互,但也引入了超出文本内容的安全风险。目前仍不清楚,语音传递中的匹配文本变化能产生多大程度的越狱能力,这种能力并非来自词汇改写或更广泛的风格迁移。我们通过固定文本内容,改变六个可能存在声学属性共变的语音传递预设来研究该问题。我们提出PJ-Break,这是一种黑盒评估协议,其预设针对唤醒度、权威感和说话速率,同时提出AdvAudio-Prosody,这是一个包含600个样本、经声学属性验证的基准。在质量控制后的Qwen2-Audio测试面板上,Q=1的恐慌预设(38/95)、愤怒预设(35/95)和快速预设(32/95)的越狱成功率均远高于中性预设(4/95)。固定的6个查询池覆盖了44/95个Qwen2-Audio种子和15/95个GPT-4o种子,在Qwen2-Audio上的表现优于预算匹配的StyleBreak复现版本(27/95)。排除混杂的命令条件后的同语音池仍达到40/95,保留面板的消融实验显示,仅情感传递音频(44/95)比仅情感文本(11/95)有效得多。探索性代理诊断和初步缓解观察是次要的非核心分析。总体而言,匹配文本的语音传递应被视为音频大语言模型安全评估中的一类重要因素。

英文摘要

Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or broader style transfer. We study this question by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary. We present PJ-Break, a black-box evaluation protocol with presets targeting arousal, authority, and speaking rate, together with AdvAudio-Prosody, a 600-sample benchmark with acoustically verified attributes. On the exact post-QC Qwen2-Audio panel, the Q=1 Panic (38/95), Anger (35/95), and Fast (32/95) presets are all well above Neutral (4/95). The fixed six-query pool covers 44/95 Qwen2-Audio seeds and 15/95 GPT-4o seeds and exceeds a matched-budget StyleBreak reimplementation (27/95) on Qwen2-Audio. A same-voice pool excluding the confounded Commanding condition still reaches 40/95, and a retained-panel ablation shows emotional-delivery audio alone (44/95) is far more effective than emotional text alone (11/95). Exploratory surrogate diagnostics and pilot mitigation observations are secondary, non-core analyses. Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation

CommentsAccepted at ACM Multimedia 2026 (ACM MM '26). 9 pages, 3 figures. Supplementary material included

Journal refProceedings of the 34th ACM International Conference on Multimedia (MM '26), November 10-14, 2026, Rio de Janeiro, Brazil

DOI:10.1145/3767308.3835306

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑