arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潜在思维链推理的结构化过程监督

Structural Process Supervision for Latent Chain-of-Thought Reasoning

Yiqi Li, Xu Chen, Chen Ju, Jiangchao Yao, Zhaoyang Li, Jinsong Lan, Xiaoyong Zhu, Yu Wang

arXiv 2609.09928首次发表:更新:

发表机构

Shanghai Jiao Tong University; Taobao & Tmall Group(上海交通大学; 淘宝天猫集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对潜在推理缺乏过程监督导致表示坍缩的问题,提出原型中介过程监督(PMPS),通过原型空间软对齐和渐进序列对齐提供结构化监督,显著压缩输出长度并提升准确率。

AI 中文摘要

潜在推理方法通过用紧凑的连续空间嵌入替代冗长的显式思维链(CoT)标记,提升了标记级效率和鲁棒性。然而,现有方法缺乏对这些潜在嵌入的直接过程监督,这常常导致表示坍缩和信息分布不均。为解决此问题,我们提出了原型中介过程监督(PMPS),该方法引入可学习的推理原型作为语义锚点,为潜在推理提供结构化的过程级监督。PMPS将潜在嵌入和显式CoT嵌入投影到共享的原型空间中,通过原型分配实现不等长表示之间的多对多软对齐。同时,我们引入了一个渐进式序列对齐(PSA)模块以进一步指导训练:位置先验最初鼓励序列对齐结构,随后逐渐放松以允许自适应匹配。实验结果表明,在GSM8K-Aug上,PMPS将输出标记长度压缩至显式CoT的50%以下。与领先基线SIM-CoT相比,我们的方法在不同模型家族中平均准确率提升了2.08%。在GPT-2上,PMPS甚至超越了CoT-SFT。在更大模型和更具挑战性的任务上,PMPS在所有具有可比输出长度的潜在推理方法中持续取得最高准确率。

英文摘要

Latent reasoning approaches enhance token-level efficiency and robustness by replacing verbose, explicit chain-of-thought (CoT) tokens with compact continuous-space embeddings. However, existing methods lack efficient process supervision over these embeddings, which often leads to representation collapse and uneven information distribution. % However, supervising these latent embeddings is challenging because a fixed number of latent states must capture information from variable-length reasoning traces. To address this, we propose Prototype-Mediated Process Supervision (PMPS), which introduces learnable reasoning prototypes as semantic anchors to provide structural process-level supervision for latent reasoning. PMPS projects latent embeddings and explicit CoT embeddings into a shared prototype space, achieving many-to-many soft alignment between unequal-length representations through prototype assignment. Meanwhile, we introduce a Progressive Sequential Alignment (PSA) module to further guide training: positional priors initially encourage sequential alignment structure, then gradually relax to permit adaptive matching. Experimental results show that PMPS compresses output token length to under 50% of explicit CoT on GSM8K-Aug. Compared to leading baseline SIM-CoT, our method achieves average accuracy gains of 2.51% across different model families. On GPT-2, accuracy of PMPS even surpasses CoT-SFT. On larger models and a more challenging task, PMPS consistently attains the highest accuracy among all latent reasoning methods with comparable output length.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑