发表机构
Shanghai Jiao Tong University; Zhejiang University(上海交通大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出UECR-GRPO,通过熵校准信用分配统一在线蒸馏与GRPO,在响应和词元级别整合验证器与教师信号,在五个数学推理基准上超越最强基线。
AI 中文摘要
基于可验证奖励的强化学习(RLVR)通过最终答案的正确性来监督数学推理,但对单个词元提供的指导很少。在线蒸馏(OPD)为学生生成的响应提供密集反馈,但教师偏好不一定反映正确性。最近的混合方法结合了OPD和验证器导出的优势,或使用教师比率重新加权任务信用。然而,教师指导在基于验证器的组归一化之后才介入,且词元重新加权不一定保持分配给每个响应的总任务信用。我们引入了用于GRPO的统一熵校准信用再分配(UECR-GRPO),它在响应和词元级别上将验证器和教师信号集成到单个GRPO风格更新中。路径效用统一(PUU)在单个KL正则化目标中结合了验证器奖励和教师到锚点路径对数比率。其在线实现使用长度归一化的教师分数,并在组归一化和PPO裁剪之前结合两种奖励,使教师证据能够影响响应排序。熵校准再分配(ECR)随后使用带符号的教师-旧策略词元差距来重新分配验证器导出的组件。全词汇教师熵衰减不确定的指导,而响应级零和投影在裁剪前保持总任务信用及其词元级符号。在五个数学推理基准上,UECR-GRPO分别使用Qwen3-1.7B和Qwen3-4B学生模型实现了17.21%和65.09%的平均Avg@12准确率,分别超过每个规模的最强基线0.89和0.56个百分点。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) supervises mathematical reasoning through final-answer correctness, but provides little guidance on individual tokens. On-policy distillation (OPD) supplies dense feedback on student-generated responses, yet teacher preference need not reflect correctness. Recent hybrids combine OPD and verifier-derived advantages or reweight task credit using teacher ratios. However, teacher guidance enters after verifier-based group normalization, and token reweighting need not preserve the total task credit assigned to each response. We introduce Unified Entropy-Calibrated Credit Redistribution for GRPO (UECR-GRPO), which integrates verifier and teacher signals within a single GRPO-style update at both the response and token levels. \emph{Path-Utility Unification} (PUU) combines verifier reward and a teacher-to-anchor path log-ratio in a single KL-regularized objective. Its on-policy implementation uses a length-normalized teacher score and combines both rewards before group normalization and PPO clipping, allowing teacher evidence to influence the response ranking. \emph{Entropy-Calibrated Redistribution} (ECR) then uses the signed teacher--old-policy token gap to redistribute the verifier-derived component. Full-vocabulary teacher entropy attenuates uncertain guidance, while a response-wise zero-sum projection preserves the total task credit and its token-wise sign before clipping. Across five mathematical reasoning benchmarks, UECR-GRPO achieves average \(\mathrm{Avg@12}\) accuracies of 17.21\% and 65.09\% with Qwen3-1.7B and Qwen3-4B students, respectively, exceeding the strongest baseline at each scale by 0.89 and 0.56 percentage points.