BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards
BV-Blend: 不确定性加权历史基线用于稳定无评论家可验证奖励强化学习
机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) ; Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China(知识驱动人机智能工程研究中心,教育部,中国) ; International Center of Future Science, Jilin University(未来科学国际中心,吉林大学)
专题命中 推理评测 :reasoning(abstract);分类 cs.AI
AI总结 针对GRPO在零方差组中学习停滞的问题,提出BV-Blend框架,通过语义聚类历史时刻与置信权重混合基线,稳定优势估计,提升训练稳定性与性能。