OPIUM:通过双目标潜在优化减轻引导外部性和过度拒绝
OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
浏览论文内容
中文总结 AI 辅助
研究激活引导向量存在的外部性问题,提出无训练方法OPIUM,通过表示匹配净化引导向量,在给定参考行为下优化新向量,改善安全与效用权衡,减轻激活引导的有害副作用。
中文摘要 AI 辅助
激活引导为在推理时控制大语言模型提供了一种轻量级机制,但引导向量可能会产生意外的外部性:效用向量可能会削弱安全行为,而拒绝向量可能会导致对良性提示的过度拒绝。我们引入了OPIUM(通过效用流形优化受保护注入),这是一种通过表示匹配来净化引导向量的无训练方法。给定两个提示集上的参考行为,OPIUM优化一个新的引导向量,该向量在保留所需干预引起的下游表示的同时,在原始向量失败的提示上匹配更安全的参考行为。在引导外部性和过度拒绝设置方面,相对于普通引导和定向消融,OPIUM改善了安全与效用的权衡,这表明激活引导的有害副作用通常可以直接在激活空间中减轻。
英文摘要
Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors may induce over-refusal on benign prompts. We introduce OPIUM (Optimizing Protected Injections via Utility Manifolds), a training-free method for sanitizing steering vectors through representation matching. Given reference behaviors on two prompt sets, OPIUM optimizes a new steering vector that preserves the downstream representations induced by the desired intervention while matching a safer reference behavior on prompts where the original vector fails. Across steering-externality and over-refusal settings, OPIUM improves the safety--utility tradeoff relative to vanilla steering and directional ablation, suggesting that harmful side effects of activation steering can often be mitigated directly in activation space.