发表机构
Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出推理时框架ALTSTEER,结合选择性干预与拒绝锚定的建设性重定向,在Llama-3.1和Qwen2.5上验证其可保持良性效用并提升建设性安全完成行为。
AI 中文摘要
安全对齐对于部署大型语言模型至关重要,要求系统在防止有害合规的同时,保持对良性请求的有用性。激活引导是一种无需训练的推理时安全控制方法,但有效的安全引导需要解决两个耦合问题:何时干预,以及干预后应如何调整生成内容。然而,现有的安全引导方法在这两个维度上仍存在局限,因为它们的触发机制在不同领域可能不稳定,且面向拒绝的引导往往会产生生硬的拒绝,而非建设性的安全指导。为解决这些局限,我们提出ALTSTEER,这是一种推理时框架,在单次推理过程中将选择性干预与以拒绝为锚点的建设性重定向相结合。ALTSTEER利用内部与拒绝相关的信号来决定何时引导,并应用分阶段引导将生成内容从面向拒绝的控制转向建设性替代方案。在Llama-3.1和Qwen2.5上的评估表明,ALTSTEER在保持良性效用的同时,提升了建设性安全完成行为,尤其是针对那些原本倾向于对有害请求生成简短拒绝的模型。
英文摘要
Safety alignment is essential for deploying large language models, requiring systems to prevent harmful compliance while preserving helpfulness on benign requests. Activation steering offers a training-free inference-time approach to safety control, but effective safety steering requires addressing two coupled questions: when to intervene and how generation should be shaped after intervention. However, existing safety steering methods remain limited along both dimensions, as their triggering mechanisms can be unstable across domains and refusal-oriented steering often yields rigid refusals rather than constructive safe guidance. To address these limitations, we propose ALTSTEER, an inference-time framework that couples selective intervention with refusal-anchored constructive redirection within a single inference pass. ALTSTEER uses an internal refusal-relevant signal to decide when to steer, and applies staged steering to shift generation from refusal-oriented control toward constructive alternatives. Evaluations on Llama-3.1 and Qwen2.5 show that ALTSTEER preserves benign utility while improving constructive safe-completion behavior, especially on models that otherwise tend to produce short refusals for harmful requests.
CommentsAccepted at EMNLP 2026 Main Conference