发表机构
Brunel University of London(布鲁内尔伦敦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究语言模型嵌入的激活引导在SFT、RLHF微调后的稳定性,发现其机制耐久但功能脆弱,需下游训练后重新验证行为。
AI 中文摘要
激活引导可直接嵌入语言模型的权重中,无需推理时干预即可塑造模型行为,还能在模型发布前编码对齐先验。但模型部署后通常会进行微调,嵌入的引导是否能留存尚不明确。本文针对5个指令微调模型(3B至14B参数)在非对抗性SFT与RLHF场景下,研究拒绝抑制和简洁性引导的嵌入引导稳定性。行为层面,留存情况与训练数据相关:当优化压力与目标行为冲突时,引导会退化,否则会留存;在SFT下,拒绝消融平均丧失64%的效果。然而机制层面,即使行为恢复,权重编辑几乎未受影响:平均向量恢复度ρ=0.004,微调沿引导方向的更新与编辑前权重模式近乎正交(平均cosθ=0.074)。当引导行为退化时,微调并非通过拆解或反转引导机制本身实现,因此嵌入引导在机制层面耐久,但功能层面脆弱,下游训练后需重新验证行为。
英文摘要
Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this. We study the stability of embedded steering for refusal ablation and brevity amplification across five instruction-tuned models (3B-14B) under non-adversarial SFT and RLHF. We find that steering degrades more under stronger optimisation pressure against the steered behaviour, which depends on both the training data and the training algorithm. Refusal ablation loses 69% of its effect on average under our SFT setup, but only 2% under our RLHF setup. The weight-edit, by contrast, remains almost unchanged (on average only 0.4% of it is restored), and the fine-tuning update along the steering direction is close to orthogonal to the pre-edit weight pattern (mean cosine similarity 0.071). Fine-tuning that restores the behaviour does so without reversing the edit. As the steered behaviour is not fully durable, embedded steering should be re-validated after downstream training.
CommentsAccepted at EMNLP 2026 Main. Earlier version at ICLR 2026 Re^4-Align Workshop