arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一阶引导:将权重适应转化为激活引导

First-Order Steering: Translating Weight Adaptation into Activation Steering

Sri Pranav Kunda, Alexander Kurz, Tomas Dominik, Uri Maoz

arXiv 2610.04283首次发表:更新:

发表机构

University of California, Davis; Chapman University; University of California, Los Angeles; Caltech(加州大学戴维斯分校; 查普曼大学; 加州大学洛杉矶分校; 加州理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一阶引导方法,将权重适应转化为激活引导向量,并开发HeRD-Merging合并过程以最小化近似误差,实现对个体及组合行为更精确的推理时控制,同时保持模型合并性能。

AI 中文摘要

激活引导利用残差流中可解释的方向,实现对模型行为的推理时操控。组合引导向量以同时应用多个目标行为,在包括AI对齐与安全在内的多个领域中十分重要,但对现有激活引导方法而言仍是一个挑战。相比之下,模型合并方面的先前工作表明,由学习到的权重适应所表示的目标行为可以高精度地组合。因此,一种将权重适应转化为激活引导向量的方法,能够扩展模型合并的先前工作,生成可组合的引导向量,从而更好地实现同时的推理时行为控制。为此,我们提出一阶引导(First-Order Steering),将激活引导表述为权重更新矩阵的一阶近似,该矩阵由引导强度向量参数化,并建立一阶引导近似误差的理论界。随后,我们开发了一种新颖的模型合并过程HeRD-Merging,其最小化一阶近似误差项,以实现更高的一阶引导精度。综合来看,我们的方法生成的引导向量在控制个体行为和组合行为方面均比现有激活引导方法更准确。此外,HeRD-Merging在匹配传统模型合并基线性能的同时,产生的权重适应能够支持更精确的一阶引导向量。

英文摘要

Activation steering exploits interpretable directions in the residual stream to enable inference-time manipulation of model behavior. Composing steering vectors to apply multiple target behaviors simultaneously is important in various fields-including AI alignment and safety-but remains a challenge for existing activation steering methods. In contrast, prior work in model merging shows that target behaviors represented by learned weight adaptations can be combined with high accuracy. A method that translates weight adaptations into activation steering vectors could therefore extend prior work in model merging to generate composable steering vectors that better enable simultaneous inference-time behavioral control. For this, we introduce First-Order Steering, a formulation of activation steering as a first-order approximation of weight update matrices parameterized by a vector of steering strengths, and establish theoretical bounds on the approximation error of first-order steering. We then develop a novel model merging procedure, HeRD-Merging, which minimizes the first-order approximation error terms to enable higher first-order steering accuracy. Together, our method produces steering vectors that control both individual and composed behaviors more accurately than existing activation steering methods. Furthermore, HeRD-Merging matches the performance of conventional model-merging baselines, while producing weight adaptations that admit more accurate first-order steering vectors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑