发表机构
School of Computer Science and Engineering, UNSW Sydney; ARC Centre of Excellence for Automated Decision-Making and Society (ADM+S); The Hong Kong University of Science and Technology (Guangzhou)(悉尼新南威尔士大学计算机科学与工程学院; 澳大利亚研究理事会自动化决策与社会卓越中心; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MetaSteer通过注意力投影适配实现上下文条件化的非线性引导,在零样本场景下达到或超越任务特定基线,并关联更强的智能体性能。
AI 中文摘要
引导大型语言模型通常依赖于在激活空间中进行线性、与上下文无关的干预,这一假设最近受到研究的质疑,并且当固定表示必须编码多种行为区分时,可能引发信息瓶颈。我们提出MetaSteer,一种学习具有上下文相关效应的非线性干预方法,并将其应用于注意力投影矩阵,从而产生随输入上下文变化而变化的激活效应,且无需线性概念几何假设。MetaSteer以偏好优化为框架,在合并的偏好语料库上训练一次,并零样本迁移至未见概念和分布外上下文。我们发现,尽管使用低秩适配器,MetaSteer仍能引发隐藏状态轨迹中结构化、上下文相关的变化,同时部分保留其局部轨迹动力学特性,包括速度和曲率。我们在三个受控文本生成基准和多个模型家族及规模的三个智能体设置中评估了MetaSteer。在零样本场景下,MetaSteer在大多数总体比较中达到或超越强任务特定引导基线。在所评估的设置中,更强的文本生成引导与更强的智能体引导性能相关联。我们进一步讨论了可迁移引导带来的几何轨迹效应、能力保留及安全性考量。
英文摘要
Steering large language models typically relies on linear, context-independent interventions in activation space, an assumption that recent work has challenged and that can induce an information bottleneck when a fixed representation must encode many behavioral distinctions. We introduce MetaSteer, a method that learns nonlinear interventions with context-dependent effects and applies them to attention projection matrices, producing activation effects that vary with the input context by construction and requiring no linear concept-geometry assumption. Framed as preference-based optimization, MetaSteer is trained once on a pooled preference corpus and transferred zero-shot to unseen concepts and out-of-distribution contexts. We find that, despite using low-rank adapters, MetaSteer induces structured, context-dependent changes in hidden-state trajectories while partially preserving aspects of their local trajectory dynamics, including velocity and curvature. We evaluate MetaSteer on three controlled text-generation benchmarks and three agentic settings across multiple model families and scales. MetaSteer matches or outperforms strong task-specific steering baselines on most aggregate comparisons in the zero-shot regime. Across the evaluated settings, stronger text-generation steering is associated with stronger agentic steering performance. We further discuss geometric trajectory effects, capability retention, and safety considerations raised by transferable steering.
CommentsPreprint. Code and pretrained model checkpoints will be released shortly