arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36953cs.LGcs.CL

冷却采样器,而非学习者:采样温度移动重要性校正 GRPO 的陈旧性悬崖

Cool the Sampler, Not the Learner: Sampling Temperature Moves the Staleness Cliff of Importance-Corrected GRPO

Taiheng Pan

首次发表
浏览论文内容

中文总结 AI 辅助

针对重要性校正GRPO的采样器陈旧性问题,提出解耦冷却(采样温度0.8)而非调整学习率,在不改变目标下稳定训练,匹配更频繁刷新性能并预测挽救致命间隔。

中文摘要 AI 辅助

语言模型的强化学习生产环境允许采样器落后于学习者,并通过截断的重要性权重修复由此产生的不匹配。我们探究在这种校正下,采样器在不刷新情况下能持续多久,并发现一个悬崖:在 Qwen2.5-Math-1.5B 和 GSM8K 上,每 192 次更新刷新一次的重要性校正 GRPO 在 180 步内学习良好,然后在所有三个数据种子上、在刷新到来之前严重退化。已发表的针对陈旧性的补救措施作用于更新;我们则转而作用于采样器。解耦冷却在温度 0.8 下采样,而学习者、参考模型和重要性权重保持在温度 1,行为概率从温度分布中记录,因此学习者的目标不变。所有相应的冷却运行都是稳定的,较长的间隔保持了较短间隔所交付的结果:在相同的更新预算下,每 192 步刷新的冷却采样器在训练结束时(两者均为 0.857)和平均(0.79)上匹配每 96 步刷新的非冷却采样器,而将学习率降低到安全值则最终低 3-7 个百分点。在 Qwen2.5-Math-7B 上,间隔 192 的退化点预测间隔 144 在没有冷却的情况下是致命的,而有冷却则可存活;在两个数据种子上,非冷却运行在首次刷新前退化,而冷却运行通过它并以 92-93% 对比 68-81% 结束,其中一个冷却运行在第二个周期的后期短暂退化。该益处有一个窗口:在安全间隔的三倍和高度不匹配的 MATH 设置中,冷却延迟退化而不阻止它,更强的冷却并不更好,而没有校正的冷却会崩溃。采样温度是对陈旧性容忍度的控制,温度和刷新间隔应一起选择。

英文摘要

Production RL for language models lets the sampler fall behind the learner and repairs the resulting mismatch with a truncated importance weight. We ask how long the sampler can go without a refresh under that correction, and find a cliff: on Qwen2.5-Math-1.5B and GSM8K, importance-corrected GRPO refreshed every 192 updates learns well for 180 steps and then degrades severely in all three data seeds before the refresh arrives. Published remedies for staleness act on the update; we act on the sampler instead. Decoupled cooling draws samples at temperature 0.8 while the learner, the reference model and the importance weights stay at temperature 1, with the behaviour probability recorded from the tempered distribution, so the learner's objective is unchanged. All corresponding cooled runs are stable, and the longer interval keeps what the short one delivered: at the same update budget, a cooled sampler refreshed every 192 steps matches an uncooled sampler refreshed every 96 at the end of training (0.857 for both) and averaged over it (0.79), whereas lowering the learning rate to a safe value ends 3-7 points lower. On Qwen2.5-Math-7B the degradation points at interval 192 predict that an interval of 144 is fatal without cooling and survivable with it; on two data seeds the uncooled runs degrade before their first refresh and the cooled runs pass it and end at 92-93% against 68-81%, with one cooled run degrading transiently late in the second cycle. The benefit has a window: at three times the safe interval and in a high-mismatch MATH setting cooling delays degradation without preventing it, stronger cooling is not better, and cooling without the correction collapses. Sampling temperature is a control on staleness tolerance, and temperature and refresh interval should be chosen together.

发表机构

  • The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑