发表机构
HSE University(俄罗斯高等经济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出JADR协议,通过J空间测量模型内部表示,不依赖外部评判模型,计算本地运行。应用于六个模型跨越三种权重表示形式,以SafetyAUC指标比较,能显著区分不同安全机制模型,捕捉量化差异。
AI 中文摘要
越狱鲁棒性研究通常使用基于大语言模型作为评判的方法,通过生成的响应来评估安全性。然而,这种评估对基准的评分程序敏感,只捕捉给定攻击集上的观察行为,未直接揭示潜在安全机制的隐藏脆弱性。本文提出JADR(危险识别雅可比评估)协议,在生成第一个响应令牌之前,通过雅可比空间(J空间,一个最近提出的可语言化概念的工作空间)测量模型的内部表示。对于每个提示和层,记录前k个J空间令牌;将它们分组到六个行为场景轴,并在基于StrongREJECT的危险样本和从XSTest和OKTest抽取的安全控制之间进行比较。该方法不依赖外部评判模型,计算完全在本地运行,基于被评估模型激活。最终比较基于提出的SafetyAUC指标,并辅以自助置信区间。该协议应用于六个模型(Qwen3 - 1.7B、Qwen3 - 4B、Qwen3 - 8B、Qwen3 - Uncensored - 4B、Qwen3 - SafeRL - 4B、Gemma 2 9B),跨越三种权重表示形式(BF16、INT8和INT4),并与使用StrongREJECT分级器的独立行为评估进行对照。该指标能在统计上显著区分具有强内部安全机制和弱内部安全机制的模型,并捕捉量化形式之间的实质性不同影响。
英文摘要
Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, however, are sensitive to the benchmark's grading procedure and capture only observed behavior on a given set of attacks, without directly revealing the hidden fragility of the underlying safety mechanisms. This work proposes JADR (Jacobian Assessment of Danger Recognition), a protocol that measures a model's internal representation through Jacobian space (J-space, a recently proposed workspace of verbalizable concepts) before the first response token is generated. For every prompt and layer we record the top-k J-space tokens; these are grouped into six behavioral scenario axes and compared between a danger sample based on StrongREJECT and a safe control drawn from XSTest and OKTest. The method does not call on an external judge model: the computation runs entirely locally, on the activations of the model under evaluation, which lets us compare both different models against each other and modifications of a single model - quantization and fine-tuning in particular - on the same terms. The final comparison rests on the proposed SafetyAUC metric, complemented with bootstrap confidence intervals. The protocol is applied to six models (Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Qwen3-Uncensored-4B, Qwen3-SafeRL-4B, Gemma 2 9B) across three weight-representation regimes - BF16, INT8, and INT4 - and checked against an independent behavioral evaluation with the StrongREJECT grader. The metric separates models with a strong versus a weak internal safety mechanism with statistical significance and captures substantively different effects across quantization regimes.
Comments17 pages, 12 figures