当技能与安全相遇:对技能合并大语言模型的自适应越狱鲁棒性进行基准测试与表征
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
浏览论文内容
中文总结 AI 辅助
本研究针对技能合并大语言模型,构建SkillSafe-Bench基准,发现静态安全无法预测其自适应越狱鲁棒性,提出SubSafe-Merge方法可在保留能力的同时消除安全侵蚀。
中文摘要 AI 辅助
模型合并已成为在不重新训练的情况下为对齐语言模型赋予新技能的默认方式:从业者通过任务算术、TIES或DARE将数学、代码或领域专家的任务向量折叠到安全对齐的基础模型中。这种便利性已知会带来安全成本,但几乎所有相关证据都基于静态拒绝测试:对固定有害提示的合规性评分。我们认为这具有误导性,因为安全对齐是“浅层的”,集中在生成的前几个标记中,合并模型的静态拒绝可能保持正常,而实际的自适应攻击仍能突破它。我们引入SkillSafe-Bench,这是一个受控基准,它采用保守的双裁判AND规则,对技能合并模型的静态拒绝、自适应越狱鲁棒性和能力保留进行评分。在六个开放权重基础模型(五个系列、两个规模)中,静态安全无法预测攻击鲁棒性:在语义模板攻击下,脆弱基础模型(Qwen的两个规模和Gemma)上看似安全的合并模型有60-76%的时间被越狱,而其他模型(Llama、Phi-4)保持鲁棒。我们进一步表明,合并的静态效应是基础模型相关的,通过无数据几何信号(任务向量与安全子空间的重叠)表征同配方的abliteration式安全侵蚀,并概述SubSafe-Merge,它通过投影去除该重叠,在保留能力的同时消除这种侵蚀。对于合并的大语言模型,自适应评估并非可选:最需要它的模型在静态筛选下看起来是安全的。
英文摘要
Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE. This convenience is known to carry a safety cost, but almost all of that evidence rests on static refusal tests: fixed harmful prompts scored for compliance. We argue this is misleading. Because safety alignment is "shallow," concentrated in the first few generated tokens, a merged model's static refusal can stay clean while a real adaptive attack still breaks it. We introduce SkillSafe-Bench, a controlled benchmark that scores skill-merged models on static refusal, adaptive jailbreak robustness, and capability retention under a conservative two-judge AND rule. Across six open-weight bases (five families, two scales), static safety does not predict robustness to attack: under a semantic template attack, safe-looking merges on the fragile bases (both Qwen scales and Gemma) are jailbroken 60-76% of the time while others (Llama, Phi-4) stay robust. We further show the static effect of merging is base-conditional, characterize same-recipe abliteration-style safety erosion through a data-free geometric signal (the overlap of a task vector with a safety subspace), and outline SubSafe-Merge, which projects this overlap away to remove that erosion at held capability. Adaptive evaluation is not optional for merged LLMs: the models that most need it look safe under static screening.
发表机构
- Google(谷歌公司)
- University of New South Wales(新南威尔士大学)
- University of Technology Sydney(悉尼科技大学)
- Zhejiang University(浙江大学)
- Australian National University(澳大利亚国立大学)
机构由 AI 辅助整理,请以论文原文为准。