发表机构
University of Technology Sydney; Southeast University; Nanjing University of Posts and Telecommunications(悉尼科技大学; 东南大学; 南京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有对抗训练生成有害行为同质化的问题,提出无目标对抗训练框架,通过放大潜在空间行为偏移生成语义多样对抗样本,以语义熵度量多样性,提升对越狱攻击的鲁棒性。
AI 中文摘要
大型语言模型(LLMs)仍然极易受到越狱攻击,这些攻击会诱导有害行为并绕过安全对齐。为了防御此类攻击,已有研究提出对抗训练范式,首先模拟失败模式,然后训练模型进行纠正,从而在安全对齐方面取得了有前景的改进。然而,这些方法通常通过鼓励固定的有害目标补全或基于固定的良性-有害数据对进行目标激活消融来构造对抗样本。因此,生成的对抗样本往往诱导同质化的有害行为,难以反映真实世界越狱攻击所引发的行为多样性。这种行为层面的狭窄性从根本上限制了其鲁棒性。为解决这一问题,我们提出了一种无目标对抗训练框架,以无监督方式生成对抗样本。通过放大并多样化模型潜在空间中的行为层面偏移,我们的方法能够产生语义多样的对抗样本,从而诱导广泛的有害行为。这种扩展的行为覆盖暴露了更多样的失败模式,进而提升了安全对齐。为量化这一效果,我们使用语义熵作为对抗行为多样性的输出层面度量。实验上,我们的方法在目标模型中引发了多样的有害行为,显著缓解了行为狭窄性,并提升了对越狱攻击的鲁棒性。
英文摘要
Large language models (LLMs) remain highly vulnerable to jailbreak attacks that induce harmful behaviors and circumvent safety alignment. To defend against such attacks, adversarial training paradigms have been proposed to first simulate failure modes and then train the model to correct them, yielding promising improvements in safety alignment. However, these methods typically construct adversarial samples either by encouraging fixed harmful target completions or by performing targeted activation ablation derived from fixed benign--harmful data pairs. As a result, the generated adversarial samples tend to induce homogeneous harmful behaviors that poorly reflect the diversity of behaviors elicited by real-world jailbreak attacks. This behavior-level narrowness fundamentally limits their robustness. To address this issue, we propose a target-free adversarial training framework that generates adversarial samples in an unsupervised manner. By amplifying and diversifying behavior-level shifts in the model's latent space, our approach produces semantically diverse adversarial samples that induce a wide range of harmful behaviors. This expanded behavioral coverage exposes more diverse failure modes and thereby improves safety alignment. To quantify this effect, we use semantic entropy as an output-level measure of adversarial behavioral diversity. Empirically, our method elicits diverse harmful behaviors in the target model, substantially mitigating behavioral narrowness and improving robustness to jailbreak attacks.