MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following
MDP-GRPO:面向多约束指令跟随的稳定化组相对策略优化
机构 * Department of Electrical and Computer Engineering, College of Engineering, University of Tehran(德黑兰大学电气与计算机工程系,工程学院) ; Department of Statistics, Mathematics and Computer Science, Allameh Tabataba’i University(塔巴蒂大学统计、数学与计算机科学系)
AI总结 针对标准GRPO在离散低分散奖励下的不稳定性,提出MDP-GRPO,通过多温度采样、双锚优势、前景理论整形和非对称KL正则化,在FollowBench等数据集上提升严格约束满足率最高5.0%。
Comments Accepted to ACL 2026 Main Conference. 14 pages, 9 figures