arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

引导向量的安全代价是可分离且可降低的

Safety Cost of Steering Vectors Is Separable and Reducible

Yuxiao Li, Gjergji Kasneci

arXiv 2608.08383首次发表:更新:

发表机构

Technical University of Munich; Munich Center for Machine Learning (MCML)(慕尼黑工业大学; 慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对引导向量破坏LLM安全机制的问题,通过识别并移除其安全退化分量,提出带约束的原始对偶优化方法,可在保留引导效果的同时大幅降低安全退化。

AI 中文摘要

引导向量是控制大语言模型(LLM)行为的轻量工具,但已有证据显示其可能意外破坏模型的安全机制,提升对有害请求的依从性,且目前尚无有效缓解方法。本研究表明,这种安全退化源于引导向量中一个可分离的分量:该分量会破坏模型安全机制,但对引导目标贡献极小。我们识别并移除了这个安全退化分量,将该任务建模为带约束的优化问题,通过原始对偶更新求解,约束条件为保留预期引导效果并限制错误弃权(不执行)。所得方案兼具可解释性与针对性:优化过程会恢复单一方向,从引导向量中移除该方向可在最小效用损失下恢复模型安全。在多个模型、引导行为及攻击套件(包括未见过的攻击类型)上,我们的方法大幅降低了引导引发的安全退化,同时保留了原始引导效果,对错误弃权的影响极小。本方法为引导向量提供了事后修正方案,缓解其安全代价,更广泛而言,它提供了一种应用激活层级模型干预的通用方案,且无需付出安全代价。

英文摘要

Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechanisms and increase compliance with harmful requests, while no effective mitigation yet exists. In this work, we show that this safety degradation arises from a separable component in the vector that disrupts the model's safety mechanisms but contributes little to the steering objective. We identify and remove this safety-degrading component, formulating the task as a constrained optimization problem solved through primal-dual updates, subject to preserving the intended steering effect and bounding false refusal. The resulting solution is both interpretable and surgical: the optimization recovers a single direction whose ablation from the steering vector restores model safety with minimal utility cost. Across models, steering behaviors, and attack suites, including unseen attacks types, our method substantially reduces steering-induced safety degradation while preserving the original steering effect with minimal impact on false refusal. Our method offers a post-hoc correction to steering vectors that mitigates their safety cost, and more broadly, it provides a general recipe for applying activation-level model interventions without paying a safety tax.

CommentsCOLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑