arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Le Critique:用于大语言模型强化学习的特权值函数

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

Siddarth Venkatraman, Matthieu Dinot, Laurence Aitchison

arXiv 2608.16739首次发表:更新:

发表机构

Mistral AI; Mila – Quebec AI Institute; Université de Montréal(米斯特拉尔人工智能公司; 米拉-魁北克人工智能研究所; 蒙特利尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对 LLM 强化学习的值函数应用难题,提出特权值函数(PVF)与自适应插值基线 TETHER,在多推理任务中提升了值函数基线性能,表现优于或媲美 GRPO。

AI 中文摘要

大语言模型(LLM)的强化学习算法主要通过方差缩减策略区分。GRPO这类组相对方法通过为每个提示采样多个 rollout 来缩减梯度方差,但仅提供序列级信用;训练还会被拖后腿的 rollout 阻碍,降低吞吐量并增加离策略程度。学习到的值函数理论上可解决这两个问题,无需大组即可提供 token 级优势,但额外的基础设施工程挑战,加上无批评者方法的实际成功,使得难以证明将其纳入 RL 流程的合理性。我们提出两种互补策略以提升值函数 RL 的性能:1)特权值函数(PVF),提供一种优雅机制注入额外的任务相关 token 级信号,且不会对策略目标产生偏差;2)TETHER,一种基线方法,根据值函数的准确性在组相对基线和值基线之间进行自适应插值。在多个推理任务中,两种策略均持续优于标准值函数基线,且与均值基线 GRPO 具有竞争力或表现更优。

英文摘要

Reinforcement learning algorithms for Large Language Models (LLMs) are largely distinguished by their variance reduction strategy. Group-relative methods like GRPO reduce gradient variance by sampling multiple rollouts per prompt, but provide only sequence-level credit. Training is also blocked by straggler rollouts, reducing throughput and increasing off-policyness. Learned value functions theoretically address both problems, providing token-level advantages without requiring large groups. However, additional infrastructure engineering challenges combined with the practical success of critic-free methods have made it difficult to justify their inclusion in RL pipelines. We propose two complementary strategies to improve the performance of value function RL: 1) Privileged Value Functions (PVF) which provide an elegant mechanism to inject additional task-relevant token-level signal without biasing the policy objective; 2) TETHER, a baseline that adaptively interpolates between group-relative and value baselines depending on the value function accuracy. Across several reasoning tasks, both strategies consistently improve over the standard value function baseline, and are competitive with or outperform mean-baseline GRPO.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑