发表机构
Alibaba Group; Tsinghua University(阿里巴巴集团; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低方差奖励下组相对优化中优势尺度校准失衡问题,提出三路校准接口及奖励分辨率协议和MaxNorm-AC方法,实现有界恢复并提升性能。
AI 中文摘要
在验证器风格的RLVR中,组相对优化通常将优势尺度视为实现细节。本文区分了两种低方差情形:不应成为偏好信号的次分辨率抖动,以及应在不扭曲KL校准的情况下学习的可信但较小的基数差距。我们提出了一种优势尺度三路校准接口:相同的组内尺度分母同时决定奖励分支强度、提示级批次权重,以及当奖励分支在原始基数尺度上重新表达时引起的有效KL校准。该接口解释了为何RLOO / this http URL 能让可信的小差距变得KL主导,而GRPO的标准差分母能无界地放大微小差距。基于该接口,我们进一步引入了奖励分辨率协议和MaxNorm-AC,分别过滤次分辨率差距并在可信非零差距上提供有界基数恢复。在密集/MoE架构和数学/代码推理上,MaxNorm-AC优于最强的鲁棒尺度基线,同时截断低方差逆尺度尾部。
英文摘要
In verifier-style RLVR, group-relative optimization often treats advantage scale as an implementation detail. This paper separates two low-variance cases: sub-resolution jitter that should not become a preference signal, and credible but small cardinal gaps that should be learned without distorting KL calibration. We propose an advantage-scale three-way calibration interface: the same within-group scale denominator simultaneously determines the reward-branch strength, prompt-level batch weight, and the effective KL calibration induced when the reward branch is re-expressed on the original cardinal scale. This interface explains why RLOO / Dr.GRPO can let credible small gaps become KL dominated, whereas GRPO's standard-deviation denominator can amplify tiny gaps without bound. Based on this interface, we further introduce the Reward-Resolution Protocol and MaxNorm-AC, respectively filtering sub-resolution gaps and providing bounded cardinal recovery on credible nonzero gaps. Across dense / MoE architectures and math / code reasoning, MaxNorm-AC improves over the strongest robust-scale baseline while truncating the low-variance inverse-scale tail.
Comments10 pages,2 figures