发表机构
Tongji University; King Abdullah University of Science and Technology(同济大学; 阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究潜在推理方法的忠实度问题,通过对输入和潜在推理步骤做特定编辑与补丁,追踪不同范式在训练检查点间忠实度演变,发现其受训练阶段和答案格式影响。
AI 中文摘要
潜在推理方法在模型连续隐藏状态中执行多步推理,有望实现更紧凑高效的推理。然而,这些不透明的隐藏状态引发了忠实度问题,即潜在推理步骤是否因果驱动最终答案。先前工作在收敛检查点研究此问题,报告了一些不忠实行为,但未考察训练期间这些行为如何形成。本文追踪不同潜在推理范式在保存检查点间忠实度的演变,对输入应用可验证反事实编辑,对潜在推理步骤应用噪声消融激活补丁。发现:在输出层面,潜在推理方法在反事实编辑下收敛时可能看似同样不忠实,但轨迹不同;在激活层面,两种范式下潜在推理步骤对最终答案的因果贡献在训练中衰减,输出翻转的示例也是贡献衰减的示例;激活层面轨迹因答案格式而异,在二元选择中衰减,在开放式解码中上升。这些发现凸显潜在推理忠实度取决于训练阶段和答案格式。
英文摘要
Latent reasoning performs multi-step inference in continuous hidden states, promising more compact and efficient reasoning. However, these opaque states raise a question of faithfulness: whether the latent reasoning steps drive the final answer. Prior work studies this question at selected checkpoints and reports several unfaithful behaviors. This endpoint view leaves how evidence of faithfulness evolves during training unexamined. We track behavioral and activation-based evidence across training using verified counterfactual edits and interventions on the latent reasoning states. We find that high task accuracy can coexist with low counterfactual responsiveness: as accuracy improves, responsiveness can decline, and different latent reasoning approaches follow distinct trajectories. On ProsQA, output sensitivity to norm-noise replacement declines alongside counterfactual responsiveness, although the result depends on the replacement. Across separately trained binary-choice and open-ended GSM settings, intervention sensitivity follows opposite trajectories. These results show that evaluating only a final checkpoint can obscure both when counterfactual responsiveness changes and what the latent states contribute.