AI 中文总结
本文提出带有动作分块的演员-评论家算法(AC2),通过评论家对动作块评分实现无需完整轨迹的信用分配,在数学证明任务上以更少计算超越GRPO。
AI 中文摘要
标准的语言模型强化学习算法将长轨迹中的每个令牌都赋予由最终奖励决定的相同优势。演员-评论家方法可以提供更细粒度的信用分配,但在使用强化学习训练大型语言模型时,学习到的评论家通常被认为过于不准确而无法信任。在最近的工作中,即使存在评论家,它也仅用于基线估计,因此每条轨迹都必须展开到其最终奖励。我们引入了带有动作分块的演员-评论家算法(AC2),该算法消除了将每条轨迹都展开到完成状态的需要。AC2 转而将信用分配给动作块:过去轨迹前缀的短延续。一个学习到的评论家对每个动作块结束时达到的状态进行评分,使策略能够在没有观察到最终奖励的情况下进行更新。我们通过三个设计选择使基于评论家的信用分配变得可靠。首先,我们引入了局部就绪性,仅在评论家对特定问题足够准确时,才在该问题上使用基于评论家的更新。其次,在可用时,我们为评论家提供来自先前成功轨迹的参考解决方案。第三,我们将信用分配给包含 10k 个令牌的动作块,而不是单个令牌,从而给评论家一个更有意义的轨迹部分进行评估。我们使用 AC2 在 FineProofs-RL 上训练 Qwen3-4B,并在 IMO-ProofBench 上进行评估。AC2 使用 2.5 倍更少的解码 FLOPs 就超过了 GRPO 的峰值验证分数 18.5%。这一收益来自两个方面:(1)AC2 达到此分数所需的训练步骤减少了 25%,(2)每一步生成的令牌更少,因为策略不需要将每条轨迹都继续到完成状态。从概念上讲,我们证明了可以消除将每条轨迹都展开到完成状态的需要,为大型语言模型强化学习算法开辟了一个巨大的、先前未探索的设计空间。
英文摘要
Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward. Actor-critic methods can provide finer-grained credit assignment, but learned critics are generally considered too inaccurate to trust when training LLMs with RL. In recent works, even when a critic is present, it is used only for baseline estimation, so every trajectory must be rolled out to its terminal reward. We introduce Actor-Critic with Action Chunking (AC2) that removes the need to roll every trajectory to completion. AC2 instead assigns credit to action chunks: short continuations of prefixes of past trajectories. A learned critic scores the state reached at the end of each action chunk, allowing the policy to update without observing a terminal reward. We make critic-based credit assignment reliable through three design choices. First, we introduce local readiness which uses critic-based updates on a problem only when the critic is sufficiently accurate on that particular problem. Second, when available, we provide the critic with a reference solution from a previous successful rollout. Third, we assign credit over action chunks of 10k tokens rather than individual tokens, giving the critic a more meaningful portion of the trajectory to evaluate. We train Qwen3-4B on FineProofs-RL using AC2 and evaluate on IMO-ProofBench. AC2 exceeds GRPO's peak validation score of 18.5% using 2.5x fewer decoding FLOPs. This gain comes from two sources, (1) AC2 requires 25% fewer training steps to reach this score, and (2) each step generates fewer tokens because the policy does not need to continue every trajectory to completion. Conceptually, we demonstrate that we can remove the need to roll out every trajectory to completion, opening up a large previously unexplored design space for LLM RL algorithms.