将潜在思维链推理解释为动力系统
Interpreting Latent CoT Reasoning as Dynamical Systems
浏览论文内容
中文总结 AI 辅助
研究针对CODI和COCONUT等潜在推理方法的可解释性问题,将潜在令牌序列建模为表示空间轨迹,用动力系统分析表征推理演变,揭示潜在思维链的动力学特性,提升其可解释性并为改进推理性能提供见解。
中文摘要 AI 辅助
近期的潜在推理方法,如CODI和COCONUT,面临基本的可解释性问题:它们在每一步在隐藏空间中维持多个叠加的候选轨迹,不像显式思维链遵循单一透明推理轨迹。现有机械方法呈现压缩、捷径和叠加,但未解释推理如何在潜在步骤中演变。为填补这一空白,我们将潜在令牌序列建模为表示空间中的轨迹,并应用动力系统分析来表征推理的演变。通过使用诸如逐步变化、方向一致性和李雅普诺夫敏感性等定量度量以及诸如UMAP和DMD/PHATE等定性投影,我们表明潜在思维链展现出具有两种不同稳定性类别的结构化、非随机动力学。CODI表现为稳定吸引子,而COCONUT表现为不稳定扩展系统,并且SIM - CoT监督在不改变潜在动力学的情况下收紧了这两种行为。该框架提升了潜在思维链推理动力学的可解释性,并为改善潜在推理性能提供了可操作的见解。代码和项目页面可在线获取。
英文摘要
Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate traces in the hidden space at each step, unlike explicit- CoT, which follows a single transparent reasoning trace. Existing mechanistic methods show compression, shortcuts, and superposition without explaining how reasoning evolves across latent steps. To address this gap, we model latent token sequences as trajectories in representation space and apply dynamical systems analysis to characterize the evolution of reasoning. Using quantitative measures, such as step-to-step change, direction consistency, and Lyapunov sensitivity, alongside qualitative projections, such as UMAP and DMD/PHATE, we show that latent CoT exhibits structured, non-random dynamics with two distinct stability classes. CODI behaves as a stable attractor, while COCONUT behaves as an unstable expanding system, and SIM-CoT supervision tightens both behaviors without changing the underlying dynamics. This framework advances the interpretability of latent CoT reasoning dynamics and provides actionable insights for improving latent reasoning performance. Code1 and Project page2 available online.
发表机构
- San Jose State University(圣何塞州立大学)
- Worcester Polytechnic Institute(伍斯特理工学院)
- George Mason University(乔治梅森大学)
- Algoverse AI Research(Algoverse人工智能研究公司)
机构由 AI 辅助整理,请以论文原文为准。