模型合并中的非对称坍塌:当拒绝覆盖识别
Asymmetric Collapse in Model Merging: When Refusal Over- writes Recognition
- UCLA(加利福尼亚大学洛杉矶分校)
- Algoverse(阿尔戈沃斯公司)
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对模型合并能否同时保留多种安全行为的问题,以Gemma-3-1B-IT微调模型为对象,发现合并时攻击抵抗性保留远多于分类准确率,原因是弃权微调的任务向量幅度更大,导致安全识别坍塌为广泛弃权。
AI中文摘要:
模型合并常被用于组合独立微调模型的能力,无需额外训练,但标准合并方法是否能同时保留多种与安全相关的行为尚不明确。我们通过受控案例研究探讨该问题,使用两个Gemma-3-1B-IT微调模型,针对两个互补安全目标:CARES伤害等级分类和WildJailbreak对抗性弃权(不执行)。我们采用Linear、SLERP、TIES和DARE-TIES四种方法合并这两个微调模型,并在分类准确率、攻击抵抗性和良性合规性上评估合并模型。在所有四种方法中,攻击抵抗性的迁移显著多于分类准确率:合并模型保留81%-85%的越狱弃权(不执行)率,而CARES准确率降至最高12.9%。权重空间测量表明,这种非对称性并非由任务向量方向强烈对立导致:两个任务向量几乎正交(余弦相似度0.011)。相反,弃权微调诱导出始终更大的每层任务向量幅度,导致对幅度敏感的方法偏向弃权更新。这些结果表明,当与安全相关的任务向量规模存在显著差异时,标准模型合并可将安全识别坍塌为广泛的弃权(不执行)行为。
英文摘要:
Model merging is often used to combine capabilities from separately fine-tuned models without additional training, but it is unclear whether standard merging methods preserve multiple safety-relevant behaviors simultaneously. We study this question through a controlled case study using two Gemma-3-1B-IT finetunes on two complementary safety objectives: CARES harm-level classification and WildJailbreak adversarial refusal. We merge the two fine-tunes using Linear, SLERP, TIES, and DARE-TIES, and evaluate the merged models on classification accuracy, attack resistance, and benign compliance. Across all four methods, attack resistance transfers significantly more than classification accuracy: merged models retain 81-85% jailbreak refusal rates while CARES accuracy falls to at most 12.9%. Weight-space measurements suggest that this asymmetry is not caused by strongly opposing task-vector directions: the two task vectors are nearly orthogonal (cosine similarity 0.011). Instead, the refusal fine-tune induces consistently larger per-layer task-vector magnitudes, causing magnitude-sensitive methods to favor refusal updates. These results show that standard model merging can collapse safety recognition into broad refusal when safety-relevant task vectors differ substantially in scale.