发表机构
New York University Abu Dhabi; New York University(纽约大学阿布扎比分校; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究审计Qwen2-Audio上AdvWave-P音频越狱的频率与解码器深度声明,发现频率排名依赖分区方式,且音频跨度发散性与频段必要性相关,但非因果定位。
AI 中文摘要
我们对Qwen2-Audio上的AdvWave-P(一种加性音频越狱)的频率和Decoder深度声明进行了审计。该协议在短时傅里叶变换(STFT)域中掩蔽扰动的频率分量,并测量攻击成功率和音频跨度表示。在520条AdvBench提示中,主要评判器将76.7%的对抗性输入标记为越狱。一项条件盲、单一标注者的验证得出该条件的Rogan-Gladen灵敏度估计为0.87(考虑验证率不确定性时约为0.83-0.95);此校正未应用于掩蔽条件。表观频率排名取决于分区方式:能量占比单独即可预测标准八频段排名(Spearman rho = 0.95),而等赫兹和等能量分区表明,掩蔽任何测试频段都能急剧降低攻击成功率。然而,在更细的16频段等能量分辨率下,掩蔽狭窄的7520-7960 Hz频段使ASR保持在0.10,若无匹配对照,此结果仍无法解决。匹配能量的分散移除测试表明,在较低频段,连续移除更具破坏性,而在较高频段,两种形式均接近下限。全局重缩放使ASR接近基线,但测试的是幅度敏感性而非频率位置。在一项重新优化试点(n = 20)中,测试的单频段和双频段支持达到0.00、0.25和0.40的ASR,而覆盖约一半STFT仓的随机支持达到0.86的平均值。在调整提示和能量的模型中,音频跨度发散性与频段必要性相关,系数从投影仪输出的+0.63上升到第30层的+0.93(对比度+0.293,95%置信区间[0.11, 0.51])。这种关联并非因果定位,单层修补也不能确立因果层。结果支持对频率声明进行分区感知审计,同时保留细分辨率顶频段结果和更广泛通用性的开放问题。
英文摘要
We audit frequency and decoder-depth claims for AdvWave-P, an additive audio jailbreak, on Qwen2-Audio. The protocol masks frequency components of the perturbation in the short-time Fourier transform (STFT) domain and measures attack success and audio-span representations. On 520 AdvBench prompts, the primary judge labels 76.7% of adversarial inputs as jailbreaks. A condition-blind, single-annotator validation yields a Rogan-Gladen sensitivity estimate of 0.87 for this condition (about 0.83-0.95 with validation-rate uncertainty); this correction is not applied to masked conditions. The apparent frequency ranking depends on the partition: energy share alone predicts the standard eight-band ranking (Spearman's rho = 0.95), and equal-Hz and equal-energy partitions show that masking any tested band can sharply reduce attack success. At a finer 16-band equal-energy resolution, however, masking the narrow 7520-7960 Hz band leaves ASR at 0.10, which remains unresolved without a matched control. Matched-energy scattered-removal tests show that contiguous removal is more damaging in lower bands, while both forms approach the floor in upper bands. Global rescaling leaves ASR near baseline but tests amplitude sensitivity rather than frequency location. In a re-optimization pilot (n = 20), tested single- and two-band supports reach ASRs of 0.00, 0.25, and 0.40, while random supports covering about half the STFT bins reach a mean of 0.86. In a prompt- and energy-adjusted model, audio-span divergence is associated with band necessity, with the coefficient rising from +0.63 at the projector output to +0.93 at layer 30 (contrast +0.293, 95% CI [0.11, 0.51]). This association is not a causal localization, and single-layer patching does not establish a causal layer. The results support partition-aware auditing of frequency claims, leaving the fine-resolution top-band result and broader generality open.