$T^5$:强化中间训练中用于令牌级思维的双评论家训练
$T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training
浏览论文内容
中文总结 AI 辅助
针对强化中间训练中令牌级信用分配难题,提出双评论家方法$T^5$,通过校准优势估计并约束漂移,提升性能7.8%,训练时间减少63.4%。
中文摘要 AI 辅助
强化中间训练使语言模型能够从未标注文本中学习内部思维,但高效的令牌级信用分配仍然具有挑战性。现有的组相对方法需要昂贵的重复生成。学习型评论家提供单次轨迹反馈,但仅凭准确的回报预测并不能确保可靠的策略更新。我们的分析表明,训练-推理不匹配和PPO裁剪阻止了优势估计中的常见偏移相互抵消,从而引入了额外的更新漂移。我们提出$T^5$,一种双评论家方法,从单次生成的轨迹中校准令牌级优势。在预热和保留资格之后,评论家提供两个优势估计,通过条件矩鞍点目标学习到的动作相关权重进行组合。该目标使每个前缀的平均优势趋向于零,而信号保留约束防止校正抹除学习信号。跨文本位置共享信息避免了重复采样每个前缀。理论上,我们刻画了信号保留约束下的最优混合,并建立了残差均值引起的漂移的上界。实验表明,与最先进的无评论家方法相比,$T^5$将平均基准性能提高了7.8%,并将平均训练步时间减少了高达63.4%。
英文摘要
Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return prediction alone does not ensure reliable policy updates. Our analysis shows how training--inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling out, introducing additional update drift. We propose \tfour{}, a twin-critic method that calibrates token-level advantages from a single generated trajectory. After warmup and held-out qualification, the critics provide two advantage estimates, combined using action-dependent weights learned through a conditional-moment saddle-point objective. This objective brings the average advantage at each prefix toward zero, while a signal-retention constraint prevents the correction from erasing the learning signal. Sharing information across text positions avoids repeated sampling of each prefix. Theoretically, we characterize optimal mixing under the signal-retention constraint and establish an upper bound on residual mean-induced drift. Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%.