通过施蒂费尔约束旋转引导控制大语言模型的拒绝行为
Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering
浏览论文内容
中文总结 AI 辅助
本研究提出基于黎曼优化的施蒂费尔约束旋转引导方法,实现高效控制大语言模型拒绝行为,经实证及消融研究验证其干预效率更优,为可靠控制LLM行为提供新方向。
中文摘要 AI 辅助
激活引导已成为推理时控制模型拒绝行为的轻量方法,越来越多研究探索可训练的激活旋转以开发几何原理的干预机制。但现有技术依赖拒绝向量等辅助结构定义这些旋转。本研究提出一种基于黎曼优化的自包含、参数高效的旋转变换学习方法,通过实证验证该方案在干预效率上更优;大量消融研究凸显了本方法关键设计选择的重要性,结果表明所提出的基于旋转的引导方案是实现更可靠控制大语言模型行为的有前景方向。
英文摘要
Activation steering has emerged as a lightweight approach for controlling model refusal at inference time. A growing line of research explores trainable rotations of activations to develop geometrically principled intervention mechanisms. However, existing techniques rely on auxiliary constructs, such as refusal vectors, to define these rotations. In our work, we develop a self-contained methodology for learning parameter-efficient rotational transformations based on Riemannian optimization. We empirically validate the proposed scheme, demonstrating its superiority in intervention efficiency. An extensive ablation study highlights the importance of key design choices in our method. Our results identify the proposed rotation-based steering scheme as a promising direction for more reliable control over the behavior of LLMs.