探测训练框架:语言模型陈旧数据强化学习比较中的设置检查
Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models
- The University of Melbourne(墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究揭示实验框架细节可逆转陈旧数据强化学习方法排名,提出PTH检查工具,验证后TIS与SAN性能持平,贡献了关键细节特征与参考结果。
中文摘要 AI 辅助
基于陈旧样本训练语言模型的方法,通过与重要性校正基线进行比较来评判。我们证明,实验框架的细节可以逆转所观察到的方法排名,并引入PTH(探测训练框架),这是一组使框架可见的检查。我们的案例是在verl和单GPU训练器上,对SAN(一种无行为方法)与截断重要性采样(TIS)进行比较,其中SAN最初在两个堆栈中均领先。框架的四个细节改变了这一比较:PPO比率是针对学习者自身重新计算的概率计算的,数据种子未到达TIS分支,重放队列在33次更新中重复使用了其第一批数据,以及两个损失归一化器与其描述不符。在每种情况下,记录的量看起来与正常工作的设置一致,而定义比较的量却未经过检查。在框架检查后,TIS在verl上与SAN持平,而在训练器中,TIS稳步学习,而SAN保持优势。我们贡献了每个细节的特征及其对比较的影响、在采样器滞后下TIS和未校正GRPO的参考结果,以及PTH检查清单。
英文摘要
Methods for training language models on stale samples are judged by comparisons against importance-corrected baselines. We show that details of the experimental harness can reverse the observed ranking of methods, and we introduce PTH (Probe The Harness), a set of checks that makes the harness visible. Our case is a comparison between SAN, a behaviour-free method, and truncated importance sampling (TIS) on verl and in a single-GPU trainer, in which SAN first finished ahead in both stacks. Four details of the harness changed this comparison: the PPO ratio was taken against the learner's own recomputed probabilities, the data seed did not reach the TIS arm, the replay queue reused its first batch for 33 updates, and two loss normalisers differed from their description. In each case the logged quantity looked consistent with a working setup, while the quantity that defines the comparison went unchecked. With the harness checked, TIS matches SAN on verl, and in the trainer TIS learns steadily while SAN keeps a margin. We contribute the signature of each detail and its effect on the comparison, reference results for TIS and uncorrected GRPO under sampler lag, and the PTH checklist.