发表机构
Sideplane AI(Sideplane AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出交换子记忆概念,通过梯度场李括号在语言模型中定位并操控路径依赖的训练历史,实现稀疏、局部读取,并能因果干预模型输出。
AI 中文摘要
不同数据上的梯度更新通常不可交换:即使使用相同的数据和总暴露量,以相反顺序在两个数据源上训练语言模型会得到不同的权重。损失或基准差异表明模型不同,但未指明差异所在。我们探究这种路径依赖性是否留下参数化的训练历史记忆:一个权重分量,当两个数据源交换时其符号翻转,在输出空间中是局部的,在定向干预下改变两种顺序之间的保留损失差距,并揭示哪个训练模型来自哪种顺序。对于在数据源A和B上各执行一次大小为$\eta$的小步SGD,权重差$\theta_{AB}-\theta_{BA}$的一阶项为$\eta^2 b_{AB}$,其中$b_{AB}=H_Bg_A-H_Ag_B$是基础模型处两个梯度场的李括号。我们通过将对数几率上的括号投影为每个词汇标记的一个分数来定义交换子记忆;这些分数之和等于括号对差距的预测。这些分数是局部的:在三个模型上,对测量的$\theta_{AB}-\theta_{BA}$或来自不相交批次的括号的相同读出,与原始前20个标记共享82-99%的重叠,而范数匹配的随机方向仅为35-49%。它们是因果可操作的:在Qwen-3-4B SFT中,对预测差距份额最大的十个标记进行降权,可关闭中位数32%的测量差距,而分数接近零的频率匹配标记几乎无效果。权重本身携带该分量:将两个训练模型之间的差异投影到$b_{AB}$上,在四个LLM中92%的情况下能识别哪个来自哪种顺序(随机概率50%)。受控测试还涵盖匹配批次DPO、冻结回放的GRPO风格目标和AdamW端点检查。该记忆按数据源对定义,而非按示例定义,其对$b_{AB}$的投影随进一步训练而衰减。
英文摘要
Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leaves a parametric training-history memory: a weight component that flips sign when the two sources are swapped, is localized in output space, changes the held-out loss gap between the two orders under targeted interventions, and reveals which trained model came from which order. For one small SGD step of size $η$ on each of sources $A$ and $B$, the weight difference $θ_{AB}-θ_{BA}$ is, to leading order, $η^2 b_{AB}$, where $b_{AB}=H_Bg_A-H_Ag_B$ is the Lie bracket of the two gradient fields at the base model. We define commutator memory by projecting the bracket through the logits into one score per vocabulary token; the scores sum to the bracket's prediction of the gap. The scores are localized: on three models, the same readout of the measured $θ_{AB}-θ_{BA}$, or of a bracket from disjoint batches, shares 82-99% of the original top-20 tokens, versus 35-49% for norm-matched random directions. They are causally actionable: in Qwen-3-4B SFT, downweighting the ten tokens with the largest predicted share of the gap closes a median 32% of the measured gap, while frequency-matched tokens with near-zero scores have almost no effect. The weights themselves carry the component: projecting the difference between the two trained models onto $b_{AB}$ identifies which came from which order in 92% of cases across four LLMs (chance 50%). Controlled tests also cover matched-batch DPO, a frozen-rollout GRPO-style objective, and an AdamW endpoint check. The memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training.
CommentsAccepted at NeurIPS 2026. 44 pages, 10 figures, 25 tables