arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10517cs.LGcs.AI

用于黎曼激活控制的条件最优桥

Conditional Optimal Bridge for Riemannian Activation Steering

Seyed Arshan Dalili, Ajay Narayanan Sridhar, Vijaykrishnan Narayanan, Mehrdad Mahdavi

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型推理时激活控制问题,提出\textsc{Cobras}方法,将其视为残差流超球面上的薛定谔桥,从优化问题推导控制目标,使控制方向查询自适应,实验显示该方法性能优于基线且避免分布外性能下降。

中文摘要 AI 辅助

激活控制为推理时控制大语言模型提供了一种轻量级的微调替代方案。许多现有方法隐式地优化期望与非期望激活分布之间的对数密度比目标,但采用启发式而非从有原则的优化问题推导。这些方法还产生与查询无关的控制方向,在分布内和分布外输入上都会降低性能。我们引入了\textsc{Cobras}(用于黎曼激活控制的条件最优桥),通过将激活控制视为残差流超球面上的薛定谔桥来解决这两个限制。这种公式化从一个适定的优化问题中得到了对数密度比控制目标的第一个有原则的推导。通过熵最优传输求解桥并提取概率流常微分方程,在Sinkhorn势均匀时可恢复广泛使用的密度比梯度。关键是,薛定谔势在当前激活处求值,使得到的控制方向本质上是查询自适应的。实验表明,在四个模型和三个对齐轴(有用性、真实性和解毒)上,\textsc{Cobras}始终优于先前的激活控制基线,同时避免了现有方法中常见的分布外性能下降。代码可在给定网址获取。

英文摘要

Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time. While many existing methods implicitly optimize a log-density-ratio objective between desired and undesired activation distributions, they do so heuristically rather than deriving it from a principled optimization problem. Moreover, these methods produce query-independent steering directions that can degrade performance on both in-distribution and out-of-distribution (OOD) inputs. We introduce \textsc{Cobras} (Conditional Optimal Bridge for Riemannian Activation Steering), which addresses both limitations by casting activation steering as a Schrödinger Bridge on the residual-stream hypersphere. This formulation yields, to our knowledge, the first principled derivation of the log-density-ratio steering objective from a well-posed optimization problem. Solving the bridge via entropic optimal transport and extracting the probability flow ODE recovers the widely used density-ratio gradient as a special case when the Sinkhorn potentials are uniform. Crucially, the Schrödinger potentials are evaluated at the current activation, making the resulting steering direction inherently query-adaptive. Empirically, across four models and three alignment axes (helpfulness, truthfulness, and detoxification), \textsc{Cobras} consistently outperforms prior activation steering baselines while avoiding the OOD degradation commonly observed in existing methods. The code can be found at https://github.com/arshandalili/cobras.

发表机构

  • The Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑