发表机构
University of Oxford; University of Utah(牛津大学; 犹他大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究推理微调对语言模型的影响,将思维链推理建模为切换动态系统,通过时间感知对比表征学习等方法恢复潜在策略。发现推理微调使模型有更丰富的潜在策略组织,恢复的状态有功能意义,SDS引导的剪枝效果良好,为推理模型分析和控制提供新视角。
AI 中文摘要
专门用于推理的语言模型相比基础模型有显著性能提升,但改善多步推理的内部变化尚不清楚。本文将思维链推理建模为切换动态系统(SDS)来解决此问题,内部表征在离散潜在策略状态下演变。通过时间感知对比表征学习和离散状态发现从激活轨迹恢复潜在策略。实验表明推理微调模型有更丰富的潜在策略组织,恢复的状态有功能专业化,因果干预证明其功能意义,SDS引导的剪枝在多数设置下优于自一致性。结果表明推理微调全局重组潜在动态,为推理模型的机制分析和过程控制提供新视角。
英文摘要
Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood. It is unclear whether reasoning fine-tuning improves local token-level competence or globally reorganizes how models structure inference over time. We address this question by modeling Chain-of-Thought reasoning as a switching dynamical system (SDS), in which internal representations evolve under discrete latent policy states. Our framework combines time-aware contrastive representation learning with discrete regime discovery to recover latent policies from activation trajectories. Across four benchmarks and model scales from 1.5B to 32B parameters, reasoning-fine-tuned models exhibit richer latent-policy organization than their base counterparts, characterized by more differentiated transition structure and model-dependent changes in state utilization, persistence, and mixing. The recovered regimes exhibit functional specialization aligned with distinct reasoning stages, and extensive controls confirm that their structure is not explained by correctness, representation learning, or modeling priors, but depends on the coherent temporal organization of reasoning trajectories. Causal interventions further show that the regimes are functionally meaningful: state-swap ablations reduce one-step predictive fit, while transplanting reasoning dynamics into base models improves performance on challenging reasoning problems. Finally, SDS-guided pruning of failure-prone reasoning prefixes outperforms self-consistency in 11 of 12 model-dataset settings, with gains of up to 12.5 percentage points. Together, our results suggest that reasoning fine-tuning globally reorganizes latent dynamics, offering a new lens for mechanistic analysis and process-level control of reasoning models.
CommentsAccepted at the Conference on Language Modeling (COLM) 2026. 45 pages, including appendices; 24 figures and 12 tables. Code: https://github.com/withmartian/mi-cot