激活流:制造用于引导的激活
Activation Flow: Manufacturing Activations for Steering
浏览论文内容
中文总结 AI 辅助
提出ActFlow,通过求解常微分方程从k个正确标签制造激活,无需微调即可引导沙袋模型,在六个锁定模型上将ARC-Easy准确率从0.05提升至0.85,接近诚实模型水平。
中文摘要 AI 辅助
差分均值引导需要记录模型表现出期望行为时的激活,而沙袋模型通过故意表现不佳来隐藏这些激活。我们引入了激活流(ActFlow),该方法无需微调即可从k个正确标签中制造这些激活。ActFlow设定目标logits,使每个标记项的正确回答排名第一,并通过在一层向所有k个残差流添加一个向量x来将logits移向这些目标。ActFlow是x的常微分方程组,每个规则对应一个将所需logit变化映射到x速度的方程。最小范数规则恰好落在目标上,而其他规则仅保留Jacobian的顶部奇异方向。我们在三个指令调优模型上测试了ActFlow,每个模型分别由沙袋提示和密码锁定的LoRA锁定。在k=40时,保留五个奇异方向的ActFlow将六个锁定模型的平均留出ARC-Easy准确率从0.05提高到0.85,而微调为0.88,诚实模型为0.92。此外,在18个锁定模型和k的组合中,它在16个组合中的得分高于最小范数规则,并且其引导方向几乎与诚实的差分均值方向正交。它还能解锁两个诚实方向失败的LoRA锁。
英文摘要
Difference-in-means steering requires activations recorded while a model shows the desired behavior, which a sandbagging model withholds by deliberately underperforming. We introduce Activation Flow (ActFlow), which manufactures these activations from $k$ correct labels without fine-tuning. ActFlow sets target logits that rank each labeled item's correct answer first, and moves the logits toward them by adding one vector $x$ to all $k$ residual streams at one layer. ActFlow is a family of ordinary differential equations for $x$, one for each rule that maps the required logit change to the velocity of $x$. The smallest-norm rule lands exactly on the targets, while the others keep only the top singular directions of the Jacobian. We test ActFlow on three instruction-tuned models, each locked by a sandbagging prompt and by a password-locked LoRA. At $k=40$, ActFlow keeping five singular directions raises the mean held-out ARC-Easy accuracy over the six locked models from $0.05$ to $0.85$, against $0.88$ for rank-$16$ LoRA fine-tuning, which trains about $3{,}000$ times as many parameters, and $0.92$ for the honest models. Furthermore, it scores higher than the smallest-norm rule in 16 of the 18 combinations of locked model and $k$, and its steering direction is nearly orthogonal to the honest difference-in-means direction. It also unlocks two LoRA locks where the honest direction fails. Surprisingly, with one labeled item, the single-step smallest-norm rule raises the Qwen2.5-7B prompt lock from $0.04$ to $0.59$.