arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20442cs.LG

存储于优化器状态,由后续训练赋予价值:阈下特质迁移的因果解释

Stored in Optimizer State, Valued by Later Training: A Causal Account of Subliminal Trait Transfer

Qinyang Xu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究揭示阈下特质迁移的两阶段机制,优化器状态传输源扰动,后续训练决定其行为价值,通过Qwen、Llama等模型实验验证了该因果机制。

中文摘要 AI 辅助

阈下特质迁移允许学生模型从教师生成的、未在语义上表达该特质的数据中获得行为倾向。近期研究解释了此类信号如何进入梯度,但未说明它们如何在源数据移除后留存,或在后续训练中获得不同符号。我们将参数和优化器矩视为单一训练器状态,推导了精确的传输-估值恒等式,将源扰动的与观察者无关的传播,与后续训练阶段及行为读出赋予的价值分离开。状态干预识别出一阶矩为因果载体:仅移植一阶矩会使参数、隐藏状态和输出在干预处保持不变,但无来源的更新会产生不断增长的参数和隐藏状态差异;将参数与一阶矩一同移植则可恢复最终行为响应。通过匹配的后续训练阶段传递相同源诱导差异,会产生Qwen效应的负、近零、正值(种子均值分别为-0.658、+0.008和+0.658)。在8次更新后,该排序在所有12个Llama-3.2-1B种子中重复出现,而不同路径的状态差异范数几乎相等;当后续训练阶段扩展至16次更新时,所有配对种子的两种对比均会增大。全视界协态可预测所有42个Qwen路径均值符号及所有21个已解析的Llama普通路径符号。与观察者无关的传输在Qwen、SmolLM2和Llama间可复现,而完整状态的重复可预测非LoRA MNIST系统中的物理、隐藏及固定头响应,包括使用AdamW和动量SGD训练的CNN。综上,这些结果确定了阈下特质迁移的两阶段机制:优化器状态传输源扰动,后续训练决定其行为价值。

英文摘要

Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically expressed. Recent work explains how such signals enter gradients, but not how they survive source removal or acquire different signs under later training. We treat parameters and optimizer moments as a single trainer state and derive an exact transport-valuation identity separating observer-independent propagation of the source perturbation from the value assigned by a future continuation and behavioral readout. State surgery identifies the first moment as a causal carrier. Transplanting it alone leaves parameters, hidden states, and outputs unchanged at the cut, yet source-free updates generate growing parameter and hidden-state differences; transplanting parameters with the first moment recovers the terminal behavioral response. Sending the same source-induced difference through matched futures produces negative, near-zero, and positive Qwen effects (-0.658, +0.008, and +0.658 seed means). This ordering recurs in all 12 Llama-3.2-1B seeds after eight updates, while state-difference norms remain nearly equal across routes. Both contrasts grow in every paired seed when the continuation extends to sixteen updates. A full-horizon costate predicts all 42 Qwen route-mean signs and all 21 resolved Llama ordinary-route signs. Observer-independent transport also replicates across Qwen, SmolLM2, and Llama, while the complete-state recurrence predicts physical, hidden, and fixed-head responses in non-LoRA MNIST systems, including CNNs trained with AdamW and momentum SGD. Together, these results identify a two-stage mechanism for subliminal trait transfer: optimizer state transports the source perturbation, and later training determines its behavioral value.

发表机构

  • Xiamen University(厦门大学)

机构由 AI 辅助整理,请以论文原文为准。

↑