arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CERO:强化学习后训练中何时何地分配回放预算

CERO: Where and When to Allocate Rollouts for RL Post-Training

Yiming Zong, Yige Wang, Xing Hu, Jiashuo Jiang, Zuo-Jun Max Shen

arXiv 2610.09679首次发表:更新:

发表机构

Hong Kong University of Science and Technology; The University of Hong Kong(香港科技大学; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CERO提出在线原始-对偶调度器,在强化学习后训练中动态分配有限回放预算,通过凹效用与Fenchel表示优化提示准入和预算节奏,在数学推理基准上取得最优平均性能。

AI 中文摘要

针对群体相对强化学习的自适应回放方法通常会在各提示词之间分配固定的每次更新预算。我们转而研究如何在完整训练时间跨度内协调有限的回放预算。我们使用累积提示暴露的凹替代效用函数来形式化该问题,并引入CERO,一种用于提示准入和预算节奏的在线原始-对偶调度器。在我们的实验中,每个被准入的提示词接收固定大小的响应组。CERO则自适应地调整选择哪些提示词、在轮次之间重新访问它们的频率,以及每轮生成多少组。紧凑的Fenchel表示将累积暴露的依赖关系线性化,而投影在线梯度下降使用奖励变化反馈和预算偏差来更新提示词特定的支撑斜率和共享预算价格。我们针对固定速率和同路径时变基准,为替代分配目标建立了路径wise保证,并显式考虑了代理差异和速率变化的项。在匹配的训练-响应预算下,CERO在三个骨干网络上的五个数学推理基准中均取得了最高的avg@16宏平均值。机制分析将CERO的提示词选择与组内奖励对比联系起来,而多种子消融实验显示,自适应节奏相比均匀和预设支出计划均带来了收益。

英文摘要

Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing. In our experiments, each admitted prompt receives a fixed-size response group. CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round. A compact Fenchel representation linearizes the dependence on cumulative exposure, while projected online gradient descent updates prompt-specific supporting slopes and a shared budget price using reward-variation feedback and budget deviations. We establish pathwise guarantees for the surrogate allocation objective against fixed-rate and same-path time-varying benchmarks, with explicit terms for proxy discrepancy and rate variation. Under matched training-response budgets, CERO attains the highest avg@16 macro-average on each of three backbones across five mathematical reasoning benchmarks. Mechanistic analyses link CERO's prompt choices to within-group reward contrast, while multi-seed ablations show gains from adaptive pacing over both uniform and preset spending schedules.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑