arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

优势尺度校准不平衡在低方差奖励下的组相对优化:诊断与有界恢复

Advantage Scale Calibration Imbalance in Group-Relative Optimization under Low-Variance Rewards: Diagnosis and Bounded Recovery

Fei Ding

arXiv 2609.19164首次发表:更新:

发表机构

Alibaba Group; Tsinghua University(阿里巴巴集团; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低方差奖励下组相对优化中优势尺度校准失衡问题,提出三路校准接口及奖励分辨率协议和MaxNorm-AC方法,实现有界恢复并提升性能。

AI 中文摘要

在验证器风格的RLVR中,组相对优化通常将优势尺度视为实现细节。本文区分了两种低方差情形:不应成为偏好信号的次分辨率抖动,以及应在不扭曲KL校准的情况下学习的可信但较小的基数差距。我们提出了一种优势尺度三路校准接口:相同的组内尺度分母同时决定奖励分支强度、提示级批次权重,以及当奖励分支在原始基数尺度上重新表达时引起的有效KL校准。该接口解释了为何RLOO / this http URL 能让可信的小差距变得KL主导,而GRPO的标准差分母能无界地放大微小差距。基于该接口,我们进一步引入了奖励分辨率协议和MaxNorm-AC,分别过滤次分辨率差距并在可信非零差距上提供有界基数恢复。在密集/MoE架构和数学/代码推理上,MaxNorm-AC优于最强的鲁棒尺度基线,同时截断低方差逆尺度尾部。

英文摘要

In verifier-style RLVR, group-relative optimization often treats advantage scale as an implementation detail. This paper separates two low-variance cases: sub-resolution jitter that should not become a preference signal, and credible but small cardinal gaps that should be learned without distorting KL calibration. We propose an advantage-scale three-way calibration interface: the same within-group scale denominator simultaneously determines the reward-branch strength, prompt-level batch weight, and the effective KL calibration induced when the reward branch is re-expressed on the original cardinal scale. This interface explains why RLOO / Dr.GRPO can let credible small gaps become KL dominated, whereas GRPO's standard-deviation denominator can amplify tiny gaps without bound. Based on this interface, we further introduce the Reward-Resolution Protocol and MaxNorm-AC, respectively filtering sub-resolution gaps and providing bounded cardinal recovery on credible nonzero gaps. Across dense / MoE architectures and math / code reasoning, MaxNorm-AC improves over the strongest robust-scale baseline while truncating the low-variance inverse-scale tail.

Comments10 pages,2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑