AI 中文总结
本研究提出可部署的实例级多层激活调控方案,解决现有全局固定层调控的缺陷,在两个8B模型和六个人格特质任务上,实现接近先验的性能,且避免流畅性崩溃。
AI 中文摘要
激活调控通过向冻结语言模型的残差流添加学习到的向量来修改模型行为,当前实践中针对每个任务全局固定注入层。我们认为最佳层的选择是实例级决策,并且我们使实例级多层选择既易于理解又可部署。在两个开放权重的8B模型和六个二元人格特质上,对层子集的实例级先验表明,最佳层因输入而异:在大多数特质-模型对上,没有固定的全局层集能恢复实例级的优势。按单层边际效应对层进行排名的贪心规则几乎能恢复先验的全部优势,但两者都需要根据标准答案对候选层进行评分,因此都无法在部署时运行;该规则转而成为仅提示预测器被训练以复刻的目标。我们的可部署方案在推理时不需要标签:一个从提示嵌入中读取的实例级层排名器、一个推断调控方向的分类器,以及一个自适应门控,该门控将短的调控后输出与推断方向进行评分,且仅调控必要数量的层。该方案恢复了先验的大部分提升(在较强模型上是主要部分,在较难模型上是明显多数),平均而言,从未使任何特质-模型对低于其未调控的对齐基线,并且在较高层数量时很大程度上避免了强全局选择会导致的流畅性崩溃。一种名为“方向而非幅度”的机制性解释,说明了在错误定向的全局集下的行为翻转、调控过多层导致的输出崩溃,以及不可调控输入的上限。
英文摘要
Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice fixes the injection layers globally per task. We argue that the best layers are an instance-level decision, and we make per-instance, multi-layer selection both well understood and deployable. On two open-weight 8B models and six binary persona traits, a per-instance oracle over layer subsets shows that the best layers vary from one input to the next: on most trait-model pairs, no fixed global layer set recovers the per-instance benefit. A greedy rule that ranks layers by single-layer marginal effect recovers nearly all of the oracle's benefit, but both must score candidate layers against the gold answer, so neither can run at deployment; the rule instead becomes the target a prompt-only predictor is trained to reproduce. Our deployable recipe needs no label at inference: a per-instance layer ranker read off the prompt embedding, a classifier that infers the steering direction, and an adaptive gate that scores short steered passes against that inferred direction and steers no more layers than necessary. The recipe recovers most of the oracle's lift (the bulk on the stronger model, a clear majority on the harder one), never drives any trait-model pair below its unsteered alignment baseline on average, and largely avoids the fluency collapse that strong global selection incurs at higher layer counts. A mechanistic account, "direction over magnitude", explains the behavioural flip under a mis-directed global set, the output collapse from steering too many layers, and the ceiling of unsteerable inputs.
Comments43 pages, 24 figures, 30 tables. Under review at ACL Rolling Review (August 2026)