一个 Token 就足够:用前缀引导桥接提示与激活引导
One Token Can Be Enough: Bridging Prompting and Activation Steering with Prefix Steering
浏览论文内容
中文总结 AI 辅助
本研究提出前缀引导(Prefix Steering)方法,通过仅对最终提示 token 后短跨度进行干预,实现与完整激活引导相当的行为控制,同时更好地保持通用能力,挑战了逐 token 引导的常见做法。
中文摘要 AI 辅助
提示通过初始上下文引导语言模型的行为,而激活引导通常在生成过程中进行干预。一个自然的问题是,引导能否对后续计算产生类似于提示的效果。在固定状态注意力假设下,我们建立了单 token 和多 token 引导匹配提示诱导的注意力头输出的充分条件,并描述了输入表示的变化如何影响这种匹配及其近似误差。这种注意力层面的联系促使我们思考,在行为层面,引导是否也能通过短暂的初始干预来指导后续生成。我们研究了前缀引导(Prefix Steering),它从最终提示 token 开始,在短跨度内应用现有的引导方向和算子,之后不再进行直接干预。我们考察了干预持续时间和强度如何共同塑造控制-能力权衡。在四个模型和五个任务中,短跨度(甚至单个 token)的干预通常能保留大部分完整引导的行为控制,同时更好地保持通用能力,提供了与提示和替代引导强度策略相比具有竞争力、在某些情况下更优的权衡。前缀引导在推理后的最终答案格式化任务上也保持有效,这表明短暂的初始干预可以影响在引导结束后很久才表达的行为。这些发现挑战了引导每个生成 token 的常见做法,并激发了对激活引导更动态的看法,即短暂干预可以在没有持续干预的情况下改变后续生成的轨迹。
英文摘要
Prompting guides language model behavior through the initial context, whereas activation steering often intervenes throughout generation. A natural question is whether steering can produce effects on subsequent computation similar to those of prompting. Under fixed-state attention assumptions, we establish sufficient conditions for single- and multi-token steering to match prompt-induced attention-head outputs, and characterize how changes in input representations affect this match and its approximation error. This attention-level connection leads us to ask whether, at the behavioral level, steering can also guide subsequent generation through a brief initial intervention. We study Prefix Steering, which applies existing steering directions and operators over a short span starting at the final prompt token, with no further direct intervention afterward. We examine how intervention duration and strength jointly shape the control-capability trade-off. Across four models and five tasks, intervention over a short span, even a single token, often retains much of full steering's behavioral control while better preserving general capabilities, offering a trade-off competitive with, and in some settings better than, prompting and alternative steering-strength policies. Prefix Steering also remains effective on final-answer formatting tasks after reasoning, suggesting that a brief initial intervention can influence behavior expressed well after steering ends. These findings challenge the common practice of steering every generated token and motivate a more dynamical view of activation steering, in which a brief intervention can alter the trajectory of subsequent generation without continued intervention.