arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

存在但重新缩放:加法激活引导的聊天到智能体的转移

Activation Steering Transfer to Agents: One Gain Ratio Does Not Identify Potency and Efficacy

Lucas Pinto

arXiv 2607.09156首次发表:更新:

AI 中文总结

研究加法激活引导从聊天到智能体的转移,通过匹配信息设计进行研究,发现转移真实但重新缩放,确定了加法特定机制,定位了重新缩放位置,指出智能体部署对引导拒绝绕过影响不可预测。

AI 中文摘要

加法激活引导(在生成过程中注入缩放后的残差流方向)几乎完全在单轮聊天中进行校准,但其目标模型越来越多地作为使用工具的ReAct智能体部署。我们首次对加法引导进行了系统的从聊天到智能体的转移研究,在匹配信息设计中结合行为测量与表征读出:相同的项目呈现为普通聊天或ReAct工具使用情节,有匹配规范的随机方向控制,并且每轮重新编码转录本以排除KV缓存污染。转移是真实的但重新缩放了,正确的描述是一种分离:在每个测试的设置和模型中(三个系列的安装站点智能体与聊天的比例为0.83 - 1.16),注入的方向以接近全强度到达后期层,而行为耦合则根据每个模型和上下文重新设置。在Qwen2.5 - 7B上,拒绝绕过向量在智能体中放大(T = 1.45,CI [\(1.20, 1.78\)],N = 300);在有动力的统一协议分布中,耦合范围从放大(Gemma - 2 - 9B,T = 2.00)到衰减(Yi - 1.5 - 9B,T = 0.43,CI [\(0.29, 0.60\)]),没有通用常数,并且有一个针对通用符号的单一干净衰减器。同一轴的定向消融不放大(T = 0.93,CI包含1),而加法注入放大(T = 1.50),增益差异为20.1个点(CI [\(13.4, 26.8\)]),这确定了一种加法特定机制。两个预先注册的工具汇聚起来,将重新缩放定位到ReAct格式支架,在任何工具观察之前,而不是到稀释理论预测的观察边界。安全影响是直接且不可预测的:智能体部署在某些模型上会将基于引导的拒绝绕过放大高达2.00倍,而其他模型则会衰减,所以部署不能假定给定模型在加法引导下是安全的。

英文摘要

Additive activation steering is calibrated in single-turn chat, then deployed inside agent scaffolds. The quantity usually reported for that move is a gain: a ratio of steered effects, T = Delta_agent / Delta_chat. We sweep eight family x arm dose-response cells over six models in both deployment contexts and show this ratio does not identify potency and efficacy. Reconstructing the published estimator in both of its forms on our own grids, its realized range contains 1 in five of five scorable cells, it moves with dose in four of five, and two cells with opposite potency shifts, both CI-clean on the primary grid, return gain intervals overlapping at a width under 0.08. Every scorable cell is an amplifier at one dose and an attenuator at another, so an amplify/attenuate taxonomy reports the dose it was read at. We replace the gain with a location: dEC50 = EC50_agent - EC50_chat, the cross-context difference in a curve location. It is signed both ways on this roster's primary grid, across four model families (+1.013 [+0.777, +1.273] against -12.368, -10.855, -5.497 and -886.066 elsewhere) and beats a vertical rescaling at equal complexity in all five cells of a frozen audit (four under the registered trim). Three deflationary accounts are measured and rejected on sign pattern and magnitude. We report the discipline at the same volume as the result: one cell is quarantined loudly, our pre-registered forecaster was refuted out of sample and is published as refuted, a registered salvage claim produced no qualifying cell and is reported unanswered, and a census of our own register reports registered branches no code here could have emitted. The consequence is a measurement instruction rather than a theorem: a transfer conclusion read at one strength does not identify what changed, because a displacement and a gain are not distinguishable from a single operating point.

Commentsv3 qualifies the two-sided sign claim to the primary dose grid: the sole positive cell loses CI-exclusion under the registered alpha*-trim. Also discloses the comparison set behind the overlapping-gain pair. 42 pages, 7 figures, 9 tables. Includes an errata section correcting v1's public record. Code: github.com/digdoug/steering-dose-reparametrization

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑