发表机构
People Make Things(People Make Things)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全双工语音对话模型的对抗攻击,提出心理声学校准的潜在平滑方法,以低质量损失显著降低劫持、静音和越狱成功率,并认证鲁棒性下限。
AI 中文摘要
端到端语音到语音对话模型同时进行听和说,因此持续开放的声学通道暴露于对抗性操纵。我们将对全双工智能体的不可感知攻击形式化为在载波语音的心理声学掩蔽阈值之下受限的加性扰动上的优化,目标有三:定向语义劫持、响应抑制和策略越狱。针对未防御的Moshi风格智能体,白盒攻击在高达91.7%的试验中成功。随后,我们引入了心理声学校准的潜在平滑(PALS),该方法在残差矢量量化潜在接口处注入由局部码本协方差塑造的各向异性高斯噪声,输入噪声由掩蔽阈值塑造以约束攻击者,并通过Kullback-Leibler一致性目标进行训练。PALS部署时无推理时间成本,在干净质量损失在2.3%以内的情况下,将劫持率降至8.3%,静音率降至11.2%,越狱率降至9.1%。蒙特卡洛平滑变体认证了高达0.616的椭球潜在半径,这是一个保证下限,而经验鲁棒性远超该下限。
英文摘要
End-to-end speech-to-speech dialogue models listen and speak simultaneously, so a continuously open acoustic channel is exposed to adversarial manipulation. We formalize imperceptible attacks on full-duplex agents as optimization over additive perturbations confined beneath the psychoacoustic masking threshold of the carrier speech, under three goals: targeted semantic hijacking, response suppression, and policy jailbreaking. Against an undefended Moshi-style agent, white-box attacks succeed in up to 91.7% of trials. We then introduce psychoacoustically aligned latent smoothing (PALS), which injects anisotropic Gaussian noise shaped by local codebook covariance at the residual-vector-quantized latent interface, with input noise shaped by the masking threshold constraining the attacker and trained by a Kullback--Leibler consistency objective. Deployed with no inference-time cost, PALS reduces hijack to 8.3%, mute to 11.2%, and jailbreak to 9.1% at clean quality within 2.3%. A Monte Carlo-smoothed variant certifies an ellipsoidal latent radius up to 0.616, a guaranteed floor that the empirical robustness far exceeds.
CommentsAccepted to IEEE SLT 2026