发表机构
Yonsei University; Korea Institute of Science and Technology (KIST); Korea University; Kyung Hee University(延世大学; 韩国科学技术研究院; 高丽大学; 庆熙大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究发现越狱鲁棒性对运行状态变化极为脆弱,仅改变普通系统提示词就可大幅调整攻击成功率,原始状态评估无法全面表征鲁棒性,需开展非原始状态下的越狱鲁棒性评估。
AI 中文摘要
现有的越狱评估通常使用默认配置(原始状态)下测量的单一攻击成功率(ASR)来表征鲁棒性。然而,用户与大语言模型(LLM)的交互会产生超出原始状态的多种运行状态。在本研究中,我们发现越狱鲁棒性对运行状态的变化极为脆弱:即使攻击保持固定,仅改变一条并非旨在影响安全性的普通系统提示词,也会显著改变攻击成功率。我们在七个对齐模型和三种代表性越狱攻击中系统研究了这一现象,观察到原始状态与非原始运行状态之间的攻击成功率存在显著差异。在某一案例中,仅因运行状态改变,攻击成功率就从2%升至58%,增幅达56个百分点。值得注意的是,这些增幅甚至出现在那些最初在原始状态评估下设计和优化的攻击中。我们进一步表明,依赖于状态的鲁棒性变化与拒斥相关轴上隐藏表征的差异存在系统性关联,且在该轴上的投影能有力预测越狱结果。我们的研究结果表明,单一的原始状态评估可能无法全面表征越狱鲁棒性,这促使人们开展评估时还需考察鲁棒性在非原始运行状态下的变化情况。
英文摘要
Existing jailbreak evaluations typically characterize robustness using a single attack success rate (ASR) measured in a default configuration (the vanilla state). However, user-LLM interactions can induce diverse operational states beyond the vanilla state. In this work, we find that jailbreak robustness is highly fragile to operational-state variation: even when the attack remains fixed, changing only an ordinary system prompt not designed to affect safety can dramatically alter attack success rates. We systematically investigate this phenomenon across seven aligned models and three representative jailbreak attacks, observing substantial differences in ASR between vanilla and non-vanilla operational states. In one case, ASR increases by up to 56 percentage points (2% to 58%) solely due to a change in operational state. Remarkably, these increases occur even for attacks originally designed and optimized under vanilla-state evaluation. We further show that state-dependent robustness variation is systematically associated with differences in hidden representations along a refusal-related axis, and that projections onto this axis strongly predict jailbreak outcomes. Our results show that a single vanilla-state evaluation may not fully characterize jailbreak robustness, motivating evaluations that also examine how robustness changes across non-vanilla operational states.
CommentsAccepted to Findings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)