跟随胜者:基于交叉熵方法的无评论家强化微调保守策略改进
Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT
浏览论文内容
中文总结 AI 辅助
针对智能体大语言模型的无评论家强化微调,提出“跟随胜者”(FTW)算法,用交叉熵方法和序数过滤替代组rollout,在Sokoban和Search-R1上匹配GRPO和PPO性能,实现方差缩减的保守策略改进。
中文摘要 AI 辅助
针对智能体大语言模型的无评论家强化微调(RFT)通常采用GRPO风格的方法,通过对重复 rollout 计算组基线来降低目标方差。然而,这种设置在作用于有状态环境(如实时服务或安全沙箱)的智能体中并不适用,因为在这些环境中重复 rollout 难以获取,且激进的更新会加剧长且稀疏验证轨迹中的噪声。我们提出“跟随胜者”(FTW),一种无评论家的策略学习算法,将交叉熵方法适配到RFT,用对回放缓冲区样本的序数过滤替代组 rollout,从而在收益的次序统计量上获得多项式集中性。我们通过控制即推理的视角推导FTW,该视角也将GRPO和DPO恢复为特定建模选择,识别出GRPO为风险中性,而DPO和FTW共享一个有界的风险追求偏移,FTW可控制该偏移。我们将此偏移识别为通过样本序数过滤进行方差缩减的固有权衡,而评论家模型则在偏差和方差之间引入不同的权衡。在智能体LLM后训练中扩展后,FTW在Sokoban和Search-R1基线上与GRPO和PPO匹配,展示了从价值模型或组rollout到CPU内存的可行权衡。
英文摘要
Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over repeated rollouts to reduce target variance. However, this setup is ill-suited to agents acting in stateful environments such as live services or security sandboxes, where repeated rollouts are impractical to obtain and aggressive updates entrench the noise of long, sparsely verified trajectories. We propose \textit{Follow the Winners} (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, replacing group rollouts with an ordinal filter on replay-buffer samples that yields polynomial concentration in the order statistic of returns. We derive FTW through a control-as-inference lens, which also recovers GRPO and DPO as specific modelling choices, identifying GRPO as risk-neutral while DPO and FTW share a bounded risk-seeking offset that FTW controls. We identify this offset as an inherent trade-off of variance reduction through ordinal filters on samples, whereas a critic model induces a different trade-off between bias and variance. Scaled to agentic LLM post-training, FTW matches GRPO and PPO on Sokoban and Search-R1 baselines, showing a viable trade-off from a value model or group rollouts to CPU memory.
发表机构
- Trent AI Limited(Trent AI有限公司)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。