arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提示主导性与不对称验证器成本:1B规模下GRPO在GSM8K上的经验消融研究

Prompt Dominance and Asymmetric Verifier Costs: Empirical Ablations of GRPO at 1B Scale on GSM8K

Yi Hou

arXiv 2610.04928首次发表:更新:

发表机构

University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文在1B规模下对GRPO进行经验消融,发现提示选择主导性能,估计器变体影响小,裁剪对离策略训练至关重要,且弱验证器在强化学习中比测试时选择更便宜。

AI 中文摘要

本文从两个方向研究1B规模下的GRPO:估计器选择对学习信号的影响,以及降级奖励信号对所学内容的影响。我们使用从头实现的GRPO在GSM8K上训练OLMo-2-0425-1B,并在受控扫描中测量两侧效果,包括一个验证器质量实验,该实验以相同方式降级训练奖励和测试时选择器。四项结果突出。提示是第一阶决策:零样本提示使基础模型停留在0.08%(其输出是退化的延续,而非错误答案),因此几乎没有组携带梯度,训练成功是因为3样本提示达到18.3%。在此规模下,估计器变体处于种子噪声范围内,Dr. GRPO在两个种子上均领先。在离策略机制中,裁剪是全部关键:在无裁剪比率的数据上训练相对于在策略参考损失4-6个百分点,而GRPO式裁剪和GSPO完全恢复损失。最后,相同的弱验证器在强化学习中比在测试时选择中便宜得多:10%翻转验证器保留强化学习的可达到增益(两个种子分别保留91%和106%),而选择仅保留57%,格式仅验证器使强化学习保留其增益的16-30%,而选择几乎一无所有。翻转噪声对期望奖励起仿射变换作用,组归一化优势与Adam的重缩放完全消除它;残余是二阶方差效应,在30%翻转率下的匹配步比较中测试。

英文摘要

This paper studies GRPO at 1B scale from both directions: what estimator choices do to the learning signal, and what a degraded reward signal does to what is learned. We train OLMo-2-0425-1B on GSM8K with a from-scratch implementation and measure both sides in controlled sweeps, including a verifier-quality experiment that degrades the training reward and the test-time selector identically. Four results stand out. The prompt is the first-order decision: the zero-shot prompt leaves the base model at 0.08% (its outputs are degenerate continuations, not wrong answers), so almost no group carries a gradient, and training succeeds because the 3-shot prompt reaches 18.3%. At this scale the estimator variants sit within seed noise, with Dr. GRPO ahead on both seeds. In the off-policy regime, clipping is the whole story: training on data without a clipped ratio loses 4-6 points relative to the on-policy reference, while GRPO-style clipping and GSPO recover the loss entirely. Finally, the same weak verifier is far cheaper in RL than in test-time selection: a 10%-flip verifier leaves RL's attainable gain intact (91% and 106% retained across two seeds) where selection retains 57%, and a format-only verifier leaves RL with 16-30% of its gain and selection with essentially nothing. Flip noise acts as an affine transform on the expected reward, and the group-normalized advantage with Adam's rescaling removes it exactly; the residual is a second-order variance effect that the matched-step comparison at a 30% flip rate tests.

Comments12 pages, 8 figures. Code, run records, and figure scripts: https://github.com/Helios-YQH/llm-from-scratch

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑