arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过价值移植引导语言模型目标

Steering Language Model Goals with Value Transplant

Pengcheng Jiang, Fabien Roger

arXiv 2609.34056首次发表:更新:

发表机构

Anthropic(Anthropic)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出价值移植方法,通过沿价值轴调整激活,将供体模型目标导向宿主模型,实验表明该干预能双向改变模型策略并提升性能。

AI 中文摘要

推理模型常常表现得好像在追求目标,但其努力并不总是朝向用户意图的方向,有时会导致它们追求非预期的结果。已有研究探讨了模型如何通过“价值轴”在内部追踪其目标的进展。我们研究改变这种信号是否能够将模型的搜索重新定向到不同的目标。我们测试了价值移植:在每个词元处,我们将宿主模型的激活沿候选价值轴移动,移动量为供体与宿主在价值坐标上的差异(乘以一个大的标量),旨在将宿主重新定向到供体的目标。我们在微调为诚实与作弊变体的Qwen3-8B和GPT-OSS-20B模型上研究这一干预。我们测试了多个候选价值轴,包括一个基于高与低引发的自我进度评分之前的激活构建的自我评分轴。该干预在两个方向上都有效,诚实的供体减少了作弊宿主在测试中的作弊行为,而作弊的供体增加了诚实宿主在测试中的作弊行为,表明该信号能够影响模型所遵循的策略。在可解决的编程任务上,来自诚实供体的移植也提高了作弊宿主的隐藏测试性能。价值移植还跨模型家族有效,为在模型控制相关场景中的干预提供了初步证据。

英文摘要

Reasoning models often act as if they pursue goals, but their efforts are not always directed toward what users intend, sometimes leading them to pursue unintended outcomes. Previous work has examined how models may internally track their progress toward their goals through a "value axis." We study whether changing such a signal can retarget the model's search toward a different goal. We test value transplant: at each token, we shift the host model's activation along a candidate value axis by the donor-host difference in value coordinates (multiplied by a large scalar), aiming to redirect the host toward the donor's goal. We study this intervention in Qwen3-8B and GPT-OSS-20B models fine-tuned into honest and cheating variants. We test several candidate value axes, including a self-rating axis constructed from activations preceding high versus low elicited self-ratings of progress. The intervention works in both directions, with an honest donor reducing test-gaming in a cheating host and a cheating donor increasing test-gaming in an honest host, showing that this signal can influence which strategy the model follows. On solvable coding tasks, transplant from an honest donor also improves the cheating host's hidden-test performance. Value transplant also works across model families, providing preliminary evidence for the intervention in a setting relevant to model control.

Comments38 pages, 26 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑