AI 中文总结
研究训练连续思维链模型,提出C-MTP直接监督方法,该方法在简化轨迹评估中表现良好,还扩展评估到复杂任务,发现当前直接和间接监督训练方法在长推理轨迹任务中表现差,揭示了现有方法的局限性。
AI 中文摘要
连续思维链方法用短序列的密集潜在表示取代冗长的推理轨迹。早期的连续思维链方法间接监督潜在表示,使其最终状态与冗长推理轨迹的状态匹配,训练时需要自回归、缓慢生成。我们引入了C-MTP,一种更简单、更快的直接监督方法,将每个潜在表示建模为要压缩的思维链轨迹中嵌入的平均值。我们的方法优于一种近似压缩令牌分布的先前直接监督方法,并且在具有简化思维链轨迹(少于100个令牌)的现有评估设置中,与较慢的间接监督方法表现相当。最后,我们将连续思维链方法的评估扩展到具有更长推理轨迹(≥数百个推理令牌)的复杂任务。我们发现,在这种情况下,直接和间接监督训练方法的表现都很差(性能下降约65%),揭示了当前连续思维链方法的局限性。代码和检查点可在该https网址获取。
英文摘要
Continuous Chain-of-Thought methods replace verbose reasoning traces with a short sequence of dense latent representations. Earlier continuous CoT methods indirectly supervise the latent representations such that its final state match that of verbose reasoning traces, requiring autoregressive, slow generation during training. We introduce C-MTP, a simpler, faster direct supervision approach that models each latent as an average of the embeddings in the CoT traces to be compressed. Our approach outperforms a prior direct supervision method that approximates the distribution of compressed tokens, and performs competitively to slower indirect supervision approaches in existing evaluation setup with simplified CoT traces (less than 100 tokens). Lastly, we extend the evaluation of Continuous CoT methods to complex tasks with longer reasoning traces ($\ge$ few hundreds reasoning tokens). We find both direct and indirect supervision training methods perform poorly (roughly 65\% performance drop) in this setting, revealing the limitations of current continuous CoT methods. The code and checkpoints are released at https://github.com/Varun221/cmtp_research
CommentsAccepted to AdaptFM Workshop, ICML 2026