arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30218cs.LGcs.AI

语言模型的微创引导

Minimally Invasive Steering of Language Models

Taha Entesari, Jingyu Zhang, Daniel Khashabi, Mahyar Fazlyab

首次发表
浏览论文内容

中文总结 AI 辅助

提出微创引导向量优化(MISVO),通过局部KL几何惩罚干预,在不更新参数的情况下优化位置特定干预,在约1B-14B参数的偏好和代码生成任务中取得最高平均奖励。

中文摘要 AI 辅助

预逻辑引导通过向冻结语言模型的最终隐藏状态添加向量,使其适应测试时奖励。无正则化的奖励优化可能显著改变输出分布并降低生成质量。我们提出了微创引导向量优化(MISVO),该方法利用诱导词元分布的局部KL几何结构来惩罚干预。由此得到的Fisher二次度量衡量分布敏感性,并通过与冻结语言模型头的矩阵-向量乘积提供解析梯度。我们推导了序列级KL梯度的精确分解,将其分为解析Fisher项和后缀得分函数项。对于固定的生成范围,我们证明后缀项在引导幅度上是二阶的,并且三个Fisher替代项在一阶上与完整KL梯度一致。MISVO使用冻结参考替代项来优化位置特定的干预,而无需更新模型参数。在参数约10亿至140亿的模型上的偏好和代码生成任务中,MISVO在七个模型-任务设置中的六个取得了最高平均奖励,其多样性和连贯性得分接近Best-of-N。

英文摘要

Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The resulting Fisher quadratic measures distributional sensitivity and admits an analytic gradient computed through matrix--vector products with the frozen language-model head. We derive an exact decomposition of the sequence-level KL gradient into an analytic Fisher term and a suffix score-function term. For a fixed generation horizon, we show that the suffix term is second order in the steering magnitude and that three Fisher surrogates agree with the full KL gradient to first order. MISVO uses the frozen-reference surrogate to optimize position-specific interventions without updating model parameters. Across preference and code-generation tasks on models with approximately 1B--14B parameters, MISVO achieves the highest mean reward in six of seven model--task settings, with diversity and coherence scores close to those of Best-of-N.

发表机构

  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑