arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TCPO:回合级信用策略优化

TCPO: Turn-Level Credit Policy Optimization

Sicong Liao, Zhi Chen, Yaohua Tang

arXiv 2608.01667首次发表:更新:

发表机构

Moore Threads AI(摩尔线程人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对验证器引导多回合强化学习的信用分配问题,提出TCPO方法,在多类任务和模型上提升了LLM的多回合推理与智能体性能。

AI 中文摘要

验证器引导的强化学习已成为提升大语言模型(LLM)推理能力的有效范式。在多回合场景中,模型每轮都会接收验证器评分并迭代优化输出。尽管这类评分提供了密集反馈,但并未直接提供密集信用:评分衡量当前输出的质量,而信用应衡量当前回合如何优化轨迹。我们提出TCPO,一种用于验证器引导多回合强化学习的回合级信用分配方法。TCPO将信用分配转化为评分到信用的转换,并通过基于参考的比较构建回合级优势:回顾性信用捕获相对于最佳先前状态的即时进展与倒退;后见之明延迟信用识别具有后续收益的无改进回合;选择性固定历史反事实估计在相同历史下优化高意外性回合。在数学推理、代码生成和AppWorld智能体任务上的实验表明,TCPO在模型规模、任务领域和验证器类型上均优于或匹配最强基线。TCPO在Qwen3-4B和DeepSeek-R1-Distill-Llama-8B上实现了最佳或并列最佳的最佳回合Pass@8,减少了成功所需的回合数,并提升了多回合智能体性能。这些结果表明,评分到信用的转换是验证器引导多回合策略优化的核心要素。

英文摘要

Verifier-guided reinforcement learning has become a powerful paradigm for improving LLM reasoning. In multi-turn settings, models receive a verifier score after each turn and iteratively refine their outputs. Although such scores provide dense feedback, they do not directly provide dense credit: a score measures the quality of the current output, while credit should measure how the current turn changes the refinement trajectory. We propose TCPO, a turn-level credit assignment method for verifier-guided multi-turn RL. TCPO casts credit assignment as score-to-credit conversion and constructs turn-level advantages through reference-based comparisons: retrospective credit captures immediate progress and regression relative to the best prior state; hindsight delayed credit identifies non-improving turns with later payoff; and selective fixed-history counterfactual estimation refines high-surprisal turns under the same history. Experiments on math reasoning, code generation, and AppWorld agent tasks show that TCPO improves or matches the strongest baselines across model scales, task domains, and verifier types. TCPO achieves the best or tied-best best-turn Pass@8 on Qwen3-4B and DeepSeek-R1-Distill-Llama-8B, reduces turns to success, and improves multi-turn agent performance. These results highlight score-to-credit conversion as a central ingredient for verifier-guided multi-turn policy optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑