PatchBench:衡量激活修补中的附带损害
PatchBench: Measuring Collateral Damage in Activation Patching
浏览论文内容
中文总结 AI 辅助
针对现有安全补丁评估忽略局部附带损害的问题,提出PatchBench基准及PatchBench-Local协议,通过生成有害与良性邻居测试补丁行为精确性,揭示聚合指标掩盖的局部回归。
中文摘要 AI 辅助
一个LLM安全补丁可能通过基准测试,但仍然是一个糟糕的修复。这种风险对于越狱修复尤为严重,其目标是纠正特定的不安全行为而不改变不相关的行为。一个补丁可能阻止精确的评估提示,却在相近的有害变体上失败,或者通过过度拒绝共享其措辞或结构的良性提示来抑制有害行为。现有协议主要测试模型是否可以被攻破,而聚合指标(攻击成功率、拒绝率、全局能力)无法区分选择性修复与更广泛的局部抑制。为解决这一差距,我们引入了PatchBench,一个基于经验观察的模型特定越狱失败基准,这些失败会诱导可操作的有害回答。从37个公共数据集的27,870个提示出发,我们筛选出15,314个英文提示,并查询了8个开源指令微调模型。结合WildGuard过滤、成对Elo排名和人工验证,我们保留了一个包含400个高置信度越狱失败的精选库。我们进一步引入了PatchBench-Local,一个测试补丁行为精确性的评估协议。对于每个有害源提示,PatchBench-Local生成三类局部邻居:保留恶意意图的有害变体、结构匹配的良性提示,以及重用关键有害术语的良性提示。它评估有害邻居的纠正和良性邻居的保留,区分选择性修复与更广泛的局部抑制。使用PatchBench-Local和MMLU评估四种激活引导方法表明,全局能力可以几乎保持不变,而局部良性回归严重,确认聚合指标遗漏了重要的附带损害。PatchBench-Local为开发和比较越狱修复方法提供了更精确的基础。
英文摘要
An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing protocols primarily test whether models can be broken, while aggregate metrics (attack success, refusal rates, global capability) cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers. Starting from 27,870 prompts from 37 public datasets, we curate 15,314 English prompts and query 8 open-source instruction-tuned models. Combining WildGuard filtering, pairwise Elo ranking, and manual verification, we retain a curated bank of 400 high-confidence jailbreak failures. We further introduce PatchBench-Local, an evaluation protocol testing whether a patch is behaviourally precise. For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants preserving malicious intent, benign prompts with matched structure, and benign prompts reusing key harmful terms. It evaluates harmful-neighbour correction and benign-neighbour preservation, distinguishing selective repair from broader local suppression. Evaluating four activation steering methods with PatchBench-Local and MMLU shows that global capability can remain nearly unchanged while local benign regressions are severe, confirming aggregate metrics miss important collateral damage. PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.
发表机构
- LIX (École Polytechnique, IP Paris, CNRS)(LIX(巴黎综合理工学院,巴黎综合理工学院联盟,法国国家科学研究中心))
- AMIAD (Agence Ministérielle pour l’IA de Défense)(AMIAD(国防人工智能部级机构))
机构由 AI 辅助整理,请以论文原文为准。