发表机构
University of Louisville(路易斯维尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估基于梯度的越狱检测在多轮对话中的脆弱性,发现其在现实良性对话上性能显著下降,需短窗口评分和长度感知阈值等校准策略才能可靠部署。
AI 中文摘要
安全对齐的语言模型通常作为多轮助手部署,这使得对抗者可以将不安全意图分散在多个用户轮次中,而非单个提示。基于梯度的越狱检测器(如GradSafe)是为单提示设计的:它们通过输入诱导梯度与固定不安全参考方向之间的对齐程度来对输入进行评分,其在多轮对话中的有效性尚不明确。我们对多轮设置下的基于梯度的越狱检测进行了受控评估。我们扩展了GradSafe,引入了一个上下文窗口扫描器,该扫描器将检测器应用于用户轮次的固定大小窗口,并使用最大窗口分数作为对话级分数。我们评估了不同的窗口大小、攻击家族、良性对话分布和目标模型。结果在合成和现实良性设置之间差异显著。针对合成良性对话,检测器在人工编写的多轮越狱上实现了0.98的ROC-AUC。在WildChat良性对话上,ROC-AUC降至0.76,且在合成数据上校准的阈值会将超过90%的良性对话标记为不安全。在现实良性分布下,单轮窗口提供最高的可分离性,而更长的窗口和累积上下文会降低性能。检测器还对攻击生成方法和目标模型敏感:成功的Crescendo攻击获得的分数与良性对话相当或更低,而Qwen2.5-7B-Instruct产生接近随机的可分离性,且具有不同的最优窗口大小。这些发现表明,基于梯度的信号可以支持多轮越狱检测,但可靠部署需要在现实良性对话上进行校准、采用短窗口评分、长度感知阈值,并跨攻击类型和模型架构进行评估。
英文摘要
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced gradient and a fixed unsafe reference direction, and their effectiveness in multi-turn dialogue remains unclear. We conduct a controlled evaluation of gradient-based jailbreak detection in multi-turn settings. We extend GradSafe with a Context Window Scanner that applies the detector to fixed-size windows of user turns and uses the maximum window score as the conversation-level score. We evaluate different window sizes, attack families, benign conversation distributions, and target models. The results differ sharply between synthetic and realistic benign settings. Against synthetic benign conversations, the detector achieves an ROC-AUC of 0.98 on human-authored multi-turn jailbreaks. On WildChat benign conversations, ROC-AUC drops to 0.76, and a threshold calibrated on synthetic data flags more than 90% of benign conversations as unsafe. Under realistic benign distributions, single-turn windows give the highest separability, whereas longer windows and accumulated contexts reduce performance. The detector is also sensitive to the attack-generation method and target model: successful Crescendo attacks receive scores comparable to or lower than benign conversations, and Qwen2.5-7B-Instruct yields near-random separability with a different optimal window size. These findings show that gradient-based signals can support multi-turn jailbreak detection, but reliable deployment requires calibration on realistic benign conversations, short-window scoring, length-aware thresholds, and evaluation across attack types and model architectures.
CommentsTo appear in CCS-LAMPS 2026