发表机构
University of Science and Technology of China; Zhejiang University(中国科学技术大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对现有SAE引导方法在复杂包装器有害提示上的不足,提出REINS方法,可抑制有害特征、增强安全弃权特征,在GUISE等数据集上大幅减少有害响应并保留模型通用能力。
AI 中文摘要
利用稀疏自编码器(SAEs)进行引导,为在不重新训练的情况下调整大型语言模型的行为提供了一种轻量级推理时路径。通过暴露稀疏且可解释的特征,SAE引导为安全控制提供了有前景的接口,可将有害的后续内容引导至弃权(不执行)。然而,我们观察到复杂的包装器仍会破坏现有SAE引导方法在有害提示上的效果。为系统评估这种失败模式,我们构建了带有复杂包装器的有害提示数据集——广义卧底指令安全评估(GUISE)。现有的单向SAE引导方法无法可靠地在有害提示上产生弃权,这表明当有害后续路径仍处于活跃状态时,仅增强弃权可能过于薄弱。这促使我们提出拒绝增强抑制引导(REINS),该方法在同一SAE特征空间中抑制有害后续特征并增强安全弃权特征。在GUISE及其他数据集上的实验表明,现有方法要么干预过弱,要么仅通过模型崩溃实现表面安全,而REINS大幅减少了有害响应,显著提升了安全弃权效果,并在很大程度上保留了模型的通用能力。
英文摘要
Steering with Sparse Autoencoders (SAEs) offers a lightweight inference-time path for adapting the behavior of large language models without retraining. By exposing sparse and interpretable features, SAE steering provides a promising interface for safety control that guides harmful continuations toward refusal. However, we observe that complex wrappers can still undermine existing SAE steering methods on harmful prompts. To evaluate this failure mode systematically, we construct Generalized Undercover Instruction Safety Evaluation (GUISE), a dataset of harmful prompts with complex wrappers. Existing single direction SAE steering methods do not reliably produce refusals on harmful prompts, suggesting that refusal enhancement alone can be too weak when the harmful continuation path remains active. This motivates us to propose Refusal-Enhanced INhibitory Steering (REINS), which suppresses harmful continuation features and enhances safe refusal features in the same SAE feature space. Experiments on GUISE and other datasets show that prior methods either intervene too weakly or achieve only apparent safety through collapse, while REINS substantially reduces harmful responses, markedly improves safe refusals and largely preserves general capabilities.
CommentsAccepted at EMNLP 2026 Main Conference