强化学习训练后计算资源应投向何处?模型大小、搜索、学习与反馈
Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback
- Technische Universität Berlin(柏林工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究强化学习训练后计算资源分配问题,引入浮点运算核算框架,通过LoRA适配的Qwen2.5策略发现条件分配前沿,揭示模型选择与训练分配关联及奖励系统对核算影响,提出RACE协议,建议相关论文报告总浮点运算及计算分配方式。
AI中文摘要:
强化学习训练后越来越多地用于使基础模型适应推理、规划和反馈驱动的机器人学习管道,但训练后资源受限常仅用单个总浮点运算预算概括。我们研究此做法背后的固定预算决策问题:在相同训练后预算下,应使用更大策略、更长时间训练较小策略、生成更多展开搜索,还是将计算用于更强奖励反馈?我们为GRPO训练后引入一个浮点运算核算框架,将计算分解为展开/搜索、策略更新/学习以及奖励或反馈模型评估。通过LoRA适配的Qwen2.5策略,我们发现条件分配前沿:最佳观察到的分配随模型大小、计算预算、奖励系统和评估目标而变化。相同浮点运算的模型大小比较表明模型选择和训练分配相互关联,因为更大策略每个令牌消耗更多计算,所以在相同预算下更新或展开次数更少。奖励系统也改变核算方式:基于规则的奖励几乎将所有非更新计算用于策略展开,而PRM风格的反馈将预算的可见部分分配给奖励模型推理。我们提出RACE作为一种诊断试点网格协议,用于在昂贵的验证运行前识别分配机制,而非保证留出改进;我们的结果表明强化学习训练后论文应报告总浮点运算以及计算在模型大小、搜索、学习和反馈之间的分配方式。
英文摘要:
Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a single total FLOP budget. We study the fixed-budget decision problem behind this practice: under the same post-training budget, should one use a larger policy, train a smaller policy longer, generate more rollout search, or spend compute on stronger reward feedback? We introduce a FLOP-accounting framework for GRPO post-training that decomposes compute into rollout/search, policy-update/learning, and reward- or feedback-model evaluation. Across LoRA-adapted Qwen2.5 policies, we find conditional allocation frontiers: the best observed allocation changes with model size, compute budget, reward system, and evaluation target. Same-FLOP model-size comparisons show that model choice and training allocation are coupled because larger policies consume more per-token compute and therefore buy fewer updates or rollouts under the same budget. Reward systems also change the accounting: rule-based rewards spend nearly all non-update compute on policy rollouts, while PRM-style feedback allocates a visible part of the budget to reward-model inference. We present RACE as a diagnostic pilot-grid protocol, not a guarantee of held-out improvement, for identifying allocation regimes before expensive validation runs; our results suggest that RL post-training papers should report total FLOPs together with how compute is divided among model size, search, learning, and feedback.