arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PACT:从信用分配到评论家对齐

PACT: From Credit Assignment to Critic Alignment

Jiayan Fu, Hang Xu, Yong Zhang, Zhaokai Luo, Yao Hu, Dongyan Zhao, Mu Chuan

arXiv 2609.26355首次发表:更新:

AI 中文总结

PACT提出token级信用分配的唯一性条件,并据此设计先演员后评论家的训练顺序,在数学推理和代码生成基准上显著超越GRPO、PPO等方法。

AI 中文摘要

强化学习已成为大型语言模型(LLM)后训练的核心组成部分,然而token级别的信用缺乏普遍接受的数学定义,导致其与常用训练信号之间的关系不明确。我们提出了三个正则性条件,即完备性、前缀一致性和中性,并证明它们唯一确定了token级别的信用。这一刻画为解释现有算法中的各种现象提供了统一基础,并指导了改进的演员-评论家训练程序的开发。通过这一视角,在策略蒸馏(OPD)中的理想教师充当隐式评论家,产生的期望策略梯度与token级别信用所诱导的梯度成正比。响应级别的REINFORCE留一法(RLOO)信号尽管粒度较粗,但与token级别信用的期望策略梯度贡献相匹配。我们进一步在奖励有界的情况下建立了近似的信用稀疏性,并展示了广义优势估计(GAE)中中间评论家误差如何变得与潜在信用相当。这些发现促使我们提出策略对齐评论家训练(PACT),它采用先演员后评论家的更新顺序,对评论家训练应用重要性采样校正,使评论家与更新后的策略更好地对齐。在智能体数学推理中,PACT在四个基准上实现了72.87%的平均准确率,分别比GRPO和PPO高出8.80和13.16个百分点。在SWE-bench Verified上,PACT达到了67.4%的通过率,分别比PPO、GRPO和SAO高出2.4、2.0和3.8个百分点。

英文摘要

Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely Completeness, Prefix Consistency, and Neutrality, and prove that they uniquely determine token-level credit. This characterization provides a unified basis for explaining phenomena across existing algorithms and guides the development of an improved actor-critic training procedure. Through this lens, an ideal teacher in On-Policy Distillation (OPD) acts as an implicit critic, yielding an expected policy gradient proportional to that induced by token-level credit. Response-level REINFORCE Leave-One-Out (RLOO) signals match the expected policy-gradient contribution of token-level credit despite their coarser granularity. We further establish approximate credit sparsity under bounded outcome rewards and show how intermediate critic errors in Generalized Advantage Estimation (GAE) can become comparable to the underlying credit. These motivate Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction to critic training and better align the critic with the updated policy. In agentic mathematical reasoning, PACT achieves 72.87% average accuracy across four benchmarks, outperforming GRPO and PPO by 8.80 and 13.16 percentage points, respectively. On SWE-bench Verified, PACT achieves a pass rate of 67.4%, outperforming PPO, GRPO, and SAO by 2.4, 2.0, and 3.8 percentage points, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑