发表机构
RWTH Aachen University; Harbin Institute of Technology, Shenzhen; The University of Tokyo; The Hong Kong University of Science and Technology(亚琛工业大学; 哈尔滨工业大学(深圳); 东京大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究区分后训练中信息被擦除、改道或缩放三种命运,提出奖励零核方法衡量信念状态变化,发现KL锚定保护信息使用,擦除仅在权重衰减下发生。
AI 中文摘要
当后训练不再奖励使用预训练模型已编码的信息时,这些信息会发生什么变化?表示压缩的常见语言混淆了三种命运:信息可能被擦除、在仍被表示的同时被改道远离决策,或在仍被表示和使用的同时被缩放以占据更少的方差。我们在预训练恢复贝叶斯信念状态的模型中使这些命运可识别。一个只读取隐藏状态粗粒度函数的奖励定义了一个精确的奖励零核。该核使我们能够分别衡量信息是否仍可恢复、决策是否因果依赖它,以及它占据多少激活方差。理论说明了什么是受保护的:KL锚定的强化学习在同等奖励的输出中保留参考策略的对数几率,监督和无锚定目标不携带此类约束,而谱压缩既不意味着擦除也不意味着使用丧失。在受控世界中,后训练主要改道或缩放奖励零信息,并使其保持可解码。没有锚定时,决策可能停止使用它尽管表示仍然存在,而有锚定时则继续使用它。擦除仅出现在长期权重衰减下,针对奖励和下一词预测都无法区分的区分。开放语言模型显示出同样的分离:上下文信念几何在晚期层谱压缩下保持可解码,类内行为取决于锚定。因此,后训练选择预训练信念状态的因果商:奖励定义决策等价性,锚定和状态更新保护它忽略的部分,优化决定其余部分是被擦除、改道还是缩放。
英文摘要
What happens to information a pretrained model already encodes when post-training no longer rewards using it? The common language of representation compression conflates three fates: information may be erased, rerouted away from the decision while still represented, or rescaled to occupy less variance while still represented and used. We make these fates identifiable in models whose pretraining recovers Bayesian belief states. A reward that reads only a coarse function of the hidden state defines an exact reward-null kernel. The kernel lets us separately measure whether the information remains recoverable, whether decisions causally depend on it, and how much activation variance it occupies. Theory says what is protected: KL-anchored reinforcement learning preserves the reference policy's log-odds among equally rewarded outputs, supervised and unanchored objectives carry no such constraint, and spectral compression implies neither erasure nor loss of use. In controlled worlds, post-training mostly reroutes or rescales reward-null information and leaves it decodable. Without an anchor decisions can stop using it although the representation survives, and with one they keep using it. Erasure appears only under prolonged weight decay, for distinctions that neither reward nor next-token prediction can see. Open language models show the same dissociation: in-context belief geometry stays decodable under late-layer spectral compression, and within-class behavior depends on the anchor. Post-training thus selects a causal quotient of the pretrained belief state: the reward defines decision-equivalence, the anchor and the state update protect part of what it ignores, and optimization decides whether the rest is erased, rerouted, or rescaled.
Comments22 pages, 16 figures, 4 tables