AI 中文总结
该研究针对潜推理的信用分配难题,提出LTC分层框架,结合多答案估计与GRPO在线策略训练,在数学推理等任务上取得最优平均准确率,可降低奖励估计误差并缓解思维级信用问题。
AI 中文摘要
潜推理使语言模型能在连续潜表示中开展中间推理,而非将其完全外化为离散思维链。然而,仅通过最终答案奖励为这类潜思维分配信用十分困难:单一最终答案会混淆思维质量与答案采样噪声。我们提出潜思维信用(Latent Thought Credit,LTC),这是一种面向潜推理的分层信用分配框架。对于每个提示,LTC会采样多个潜思维,固定每个思维后的上下文,并通过对该固定上下文生成的多个答案的奖励取平均,来估计思维级期望奖励。LTC利用思维级优势优化潜思维阶段,利用答案级优势优化答案阶段,并采用优势加权的思维匹配目标,以帮助策略复现高信用潜思维。我们将LTC实例化为GRPO风格的在线策略训练框架,并在数学推理和STEM多项选择题任务上进行评估。LTC在对比方法中取得了最佳平均准确率,而消融实验和固定上下文诊断表明,多答案估计可降低奖励估计误差,并缓解模糊或不正确的思维级信用问题。
英文摘要
Latent reasoning allows language models to carry out intermediate reasoning in continuous latent representations rather than fully externalizing it as discrete chains of thought. However, assigning credit to such latent thoughts from answer-only rewards is difficult: a single final answer mixes thought quality with answer-sampling noise. We propose \textbf{Latent Thought Credit (LTC)}, a hierarchical credit-assignment framework for latent reasoning. For each prompt, LTC samples multiple latent thoughts, fixes the context after each thought, and estimates thought-level expected reward by averaging rewards over multiple answers generated from that fixed context. LTC uses thought-level advantages to optimize the latent-thought phase, answer-level advantages to optimize the answer phase, and an advantage-weighted thought-matching objective that helps the policy reproduce high-credit latent thoughts. We instantiate LTC in a GRPO-style on-policy training framework and evaluate it across mathematical reasoning and STEM multiple-choice tasks. LTC achieves the best average accuracy among the compared methods, while ablations and fixed-context diagnostics show that multi-answer estimation reduces reward-estimation error and mitigates ambiguous or incorrect thought-level credit.