发表机构
University of Michigan; University of Waterloo(密歇根大学; 滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究审计 GRPO、SFT、DPO 三种后训练方法对语言模型上下文 grounding 的提升,发现其提升大多依赖初始模型已有的机制,而非新增机制。
AI 中文摘要
语言模型会在提示证据与记忆知识冲突时忽略该证据。后训练可使模型更可靠地遵循此类证据,但尚不清楚这些提升是否需要新机制,还是强化已存在的机制。我们从同一初始检查点出发,对比了 GRPO、SFT 和 DPO 共九种后训练方案,关键对比跨模型规模与家族扩展。我们估计了训练前该检查点的 grounding 方向。在测试的五种 GRPO 变体中,grounding 提升幅度很小;对跨种子重复的两种变体,等价性检验将其效应约束在冲突-SFT 提升以下,即便奖励指标有所改善。冲突-SFT 使 grounding 适度提升,而 DPO 在其匹配分布上使 grounding 接近天花板。冲突-SFT 和 DPO 大多使用与初始模型相同的因果注意力头集。减去初始模型方向会抑制两种提升,而将其添加到初始模型可恢复 DPO 35% 的提升,且剂量通过所有所述副作用检查。在监督预热使上下文答案出现在更多 rollout 后,相同 GRPO 方案基本无进一步 grounding 提升。在我们的设定中,grounding 提升大多依赖初始模型已存在的机制。
英文摘要
Language models can ignore prompt evidence when it conflicts with memorized knowledge. Post-training can make models follow such evidence more reliably, but it is unclear whether these gains require new machinery or strengthen machinery already present. We compare nine post-training arms spanning GRPO, SFT, and DPO from one starting checkpoint, with key comparisons extended across scales and families. We estimate a grounding direction from that checkpoint before training. Across five tested GRPO variants, grounding gains are small. For the two variants replicated across seeds, equivalence tests bound their effects below the conflict-SFT gain even as the rewarded metric improves. Conflict-SFT improves grounding moderately, while DPO drives grounding near ceiling on its matched distribution. Conflict-SFT and DPO largely use the same causal attention-head set as the starting model. Subtracting the starting-model direction suppresses both gains, while adding it to the starting model recovers 35% of DPO's gain at a dose passing all stated side-effect checks. After a supervised warm start makes the context answer appear in more rollouts, the same GRPO recipe adds essentially no further grounding gain. In our setting, grounding gains largely depend on machinery already present in the starting model.