发表机构
King Abdullah University of Science and Technology; Stony Brook University(阿卜杜拉国王科技大学; 石溪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有LRMs安全对齐方法依赖模式泛化不足的问题,提出双对抗框架AdvSafe,通过两阶段对抗博弈生成不安全知识数据集,使LRMs实现鲁棒安全对齐且权衡性能更优。
AI 中文摘要
大型推理模型(LRMs)在复杂任务上取得了显著成功,但仍易受诱导产生不安全输出的有害提示影响。近期方法通过直接拒绝或安全理由对齐LRMs,却往往聚焦于提示模式而非内在攻击机制。因此,这些以模式为中心的对齐方法在不同越狱攻击中难以泛化,损害了对抗鲁棒性和推理实用性。我们提出AdvSafe,一种双对抗框架,通过明确解构对抗机制使LRMs内化不安全知识。这突破了依赖模式的痕迹,在不损害推理实用性的前提下培养了鲁棒认知防御。我们的流程通过两阶段对抗博弈运行:第一阶段为对抗合成,自主智能体动态制作欺骗性越狱提示,调整策略以突破强教师模型;第二阶段为对抗提取,被突破的教师执行认知反击。对于每一次成功的越狱,教师会揭露伪装,解释攻击为何成功以及如何识别和缓解此类提示。这一双对抗过程生成了包含丰富、可泛化不安全知识的紧凑推理数据集,基于此数据集训练的学生模型通过内在威胁理解隐含地获得安全对齐。实验表明,仅需1000个合成样本,经AdvSafe对齐的LRMs相比现有基线实现了显著更强的越狱鲁棒性,且几乎无实用性下降;此外,AdvSafe提升了对分布外提示的鲁棒性,证明学习不安全知识可实现更优的鲁棒性-实用性权衡,并能泛化至未见过的攻击模式。
英文摘要
Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent methods align LRMs using direct refusals or safety rationales, yet often focus on prompt patterns rather than intrinsic attack mechanisms. As a result, these pattern-centric alignments struggle to generalize across diverse jailbreaks, compromising adversarial robustness and reasoning utility. We propose AdvSafe, a dual-adversarial framework that enables LRMs to internalize unsafety knowledge by explicitly deconstructing adversarial mechanisms. This moves beyond pattern-dependent traces, fostering robust cognitive defense without compromising reasoning utility. Our pipeline operates via a two-phase adversarial game. First, in adversarial synthesis, an autonomous agent dynamically crafts deceptive jailbreak prompts, adapting its strategies to breach a strong teacher model. Second, in adversarial extraction, the breached teacher executes a cognitive counter-attack. For every successful jailbreak, the teacher unmasks the camouflage, explaining why the attack succeeds and how such prompts can be identified and mitigated. This dual-adversarial process yields a compact reasoning dataset capturing rich, generalizable unsafety knowledge. Student models trained on this dataset implicitly acquire safety alignment through intrinsic threat comprehension. Experiments show that with only 1K synthesized samples, AdvSafe-aligned LRMs achieve significantly stronger jailbreak robustness than existing baselines, with almost no utility degradation. Furthermore, AdvSafe improves robustness against out-of-distribution prompts, demonstrating that learning unsafety knowledge enables a superior robustness-utility trade-off and generalizes beyond seen attack patterns.