arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32791cs.AI

$T^5$:强化中间训练中用于令牌级思维的双评论家训练

$T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training

Nan Qiao, Yebin Yang, Weinong Wang, Shuning Wang, Shangpin Peng, Fengyuan Lu, Xinming Wang, Zhehan Kan, Ruixu Zhang, Songyang Zhang, Sheng Yue, Yonglong Tian, Ju Ren

首次发表
浏览论文内容

中文总结 AI 辅助

针对强化中间训练中令牌级信用分配难题,提出双评论家方法$T^5$,通过校准优势估计并约束漂移,提升性能7.8%,训练时间减少63.4%。

中文摘要 AI 辅助

强化中间训练使语言模型能够从未标注文本中学习内部思维,但高效的令牌级信用分配仍然具有挑战性。现有的组相对方法需要昂贵的重复生成。学习型评论家提供单次轨迹反馈,但仅凭准确的回报预测并不能确保可靠的策略更新。我们的分析表明,训练-推理不匹配和PPO裁剪阻止了优势估计中的常见偏移相互抵消,从而引入了额外的更新漂移。我们提出$T^5$,一种双评论家方法,从单次生成的轨迹中校准令牌级优势。在预热和保留资格之后,评论家提供两个优势估计,通过条件矩鞍点目标学习到的动作相关权重进行组合。该目标使每个前缀的平均优势趋向于零,而信号保留约束防止校正抹除学习信号。跨文本位置共享信息避免了重复采样每个前缀。理论上,我们刻画了信号保留约束下的最优混合,并建立了残差均值引起的漂移的上界。实验表明,与最先进的无评论家方法相比,$T^5$将平均基准性能提高了7.8%,并将平均训练步时间减少了高达63.4%。

英文摘要

Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return prediction alone does not ensure reliable policy updates. Our analysis shows how training--inference mismatch and PPO clipping prevent a common offset in advantage estimates from cancelling out, introducing additional update drift. We propose \tfour{}, a twin-critic method that calibrates token-level advantages from a single generated trajectory. After warmup and held-out qualification, the critics provide two advantage estimates, combined using action-dependent weights learned through a conditional-moment saddle-point objective. This objective brings the average advantage at each prefix toward zero, while a signal-retention constraint prevents the correction from erasing the learning signal. Sharing information across text positions avoids repeated sampling of each prefix. Theoretically, we characterize optimal mixing under the signal-retention constraint and establish an upper bound on residual mean-induced drift. Experiments show that, compared with the state-of-the-art critic-free method, \tfour{} improves mean benchmark performance by 7.8\% and reduces mean training-step time by up to 63.4\%.

↑