AI 中文总结
该研究针对大语言模型越狱后缀优化中候选选择的短视问题,提出轨迹感知候选选择框架\textit{OURS},在HarmBench上优于基线,提升了攻击成功率与优化稳定性。
AI 中文摘要
基于梯度的越狱后缀优化方法通常通过保留当前损失最低的候选来更新后缀。我们发现,这一看似自然的设计存在根本性的短视问题:在当前步骤代理下表现更好的候选,往往在后续搜索中无法产生更好的越狱结果,这揭示了一种选择阶段的奖励黑客行为。这表明,候选选择而非仅候选生成,是后缀优化中一个隐藏的瓶颈。为解决该问题,我们提出\textit{OURS},一种面向越狱后缀优化的轨迹感知候选选择框架。\textit{OURS}不再仅通过即时损失选择候选,而是用轨迹感知代理增强每一步评估,并通过参考策略正则化和判别器估计的卡方校正来稳定选择,鼓励在当前步骤之外仍有效的选择。在HarmBench上的实验表明,在相同搜索预算下,\textit{OURS}始终优于强大的基线,大幅提高了攻击成功率,同时在整个搜索过程中表现出更稳定的优化行为。我们的发现强调,缓解短视候选选择导致的选择阶段奖励黑客行为,对改进越狱后缀优化至关重要。
英文摘要
Gradient-based jailbreak suffix optimization methods typically update the suffix by retaining the candidate with the lowest current loss. We show that this seemingly natural design is fundamentally myopic: candidates that look better under the current-step proxy often fail to produce better jailbreak outcomes later in the search, revealing a form of selection-stage reward hacking. This suggests that candidate selection, rather than candidate generation alone, is a hidden bottleneck in suffix optimization. To address this issue, we propose TACS, a trajectory-aware candidate selection framework for jailbreak suffix optimization. Instead of selecting candidates solely by their immediate loss, TACS augments per-step evaluation with a trajectory-aware proxy and stabilizes selection with reference-policy regularization and a discriminator-estimated chi-squared correction, encouraging choices that remain effective beyond the current step. Experiments on HarmBench show that TACS consistently outperforms strong baselines under the same search budget, substantially improving attack success rates while exhibiting more stable optimization behavior throughout the search. Our findings highlight that mitigating selection-stage reward hacking caused by myopic candidate selection is critical for improving jailbreak suffix optimization.
CommentsWe identified an error in the theoretical analysis, which affects the validity of the main conclusions of the manuscript. Since the current version does not adequately support this conclusion, we have decided to withdraw the paper