arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

免训练任务向量用于大语言模型行为控制

Training-Free Task Vectors for LLM Behavioral Control

Gabriel J. Perin, Lucas Boscaini, André Araujo, Nina S. T. Hirata

arXiv 2609.09054首次发表:更新:

发表机构

University of São Paulo; Google; Google DeepMind(圣保罗大学; 谷歌; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出免训练任务向量(TFTVs),无需微调即可通过前向传播统计量计算任务向量,实现大语言模型行为的放大、抑制与组合,且保持通用能力。

AI 中文摘要

任务向量通过识别权重空间中语义上有意义的方向来实现训练后的模型编辑,通常计算为微调模型与其预训练初始化之间的差异。然而,这种对微调过程的依赖使得发现此类方向成本高昂,并限制了训练后模型编辑的实用性。为解决这一局限,我们提出了免训练任务向量(TFTVs),一种无需微调即可计算类任务向量方向的新方法。我们的方法仅利用前向传播统计量将激活引导向量映射为秩一权重空间编辑,同时满足算术性质,直接支持通过加法进行学习、通过减法进行遗忘以及多个编辑的组合。在实验上,我们在大型语言模型行为控制任务上评估了TFTVs,并表明它们能一致地放大、抑制和组合目标行为,同时保持通用知识和问题解决能力。我们还与其他编辑和引导基线方法进行了验证,实验证明TFTVs在实现更强特质控制的同时,具有更好或具有竞争力的效用保持。我们希望我们的工作为社区在训练后模型编辑和更广泛的免训练模型控制方面开辟新方向。代码可在项目网站获取:此HTTP URL。

英文摘要

Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality of post-training model editing. To address this limitation, we introduce Training-Free Task Vectors (TFTVs), a novel method to compute task-vector-like directions without requiring fine-tuning. Our method maps activation steering vectors to rank-one weight-space edits using only forward-pass statistics, while satisfying arithmetic properties that directly support learning via addition, forgetting via subtraction, and the composition of multiple edits. Empirically, we evaluate TFTVs on large language model behavioral control tasks and show that they consistently amplify, suppress, and compose target behaviors while preserving general knowledge and problem-solving skills. We also validate our method against other editing and steering baselines, experimentally demonstrating that TFTVs achieve stronger trait control with better or competitive utility preservation. We hope our work opens new directions for the community in post-training model editing and broader training-free model control. Code is available on the project website: tftv-llm.github.io.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑