arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

后训练在无关决策上留下行为阴影

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Ziyang Zhang, Yubin Jing, Yuanhao Zeng, Yuyao Li, Haofan Wang, Yichen Gong

arXiv 2609.29233首次发表:更新:

发表机构

Peking University; Georgia Institute of Technology; ShanghaiTech University; Tsinghua University; Lovart AI(北京大学; 佐治亚理工学院; 上海科技大学; 清华大学; Lovart AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出主动无任务蒸馏(ATD),利用教师模型单词语义差异探测后训练行为阴影,实现跨无关文本的能力迁移,并在编码等任务上验证了显著效果。

AI 中文摘要

我们发现语言模型可以通过与任务无关的文本迁移能力。后训练通常使用特定任务的数据来改进语言模型。先前关于潜意识学习的研究表明,这些更新的信息可以通过无关的生成内容传递,但主要集中于使用大量教师输出的特质或偏好。我们引入了主动无任务蒸馏(ATD),该方法仅使用教师每个提示中的一个单词即可实现能力迁移。ATD通过选择教师和学生共享的公共祖先在两个普通单词之间几乎无差异的提示来探测后训练的行为阴影。从该祖先初始化的学生仅从由此产生的提示-单词对中学习,无需目标任务示例、教师逻辑或教师参数。在主要编码实验中,使用Qwen2.5-1.5B,5,664个提示在HumanEval+上相对于精确干扰匹配的对照(该对照破坏提示-响应关联)获得了5.34个百分点的提升。进一步实验显示,在额外的模型代际、规模和家族中,科学知识、常识推理和阅读理解方面的迁移。功能分析表明,学习到的阴影是可组合的,且其强度跟踪教师的更新强度。

英文摘要

We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.

Comments17 pages, 6 figures, 13 tables. Code: https://github.com/myboker/ATD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑