STAR-GRPO:针对表示依赖的奖励黑客的规范锚定与可靠性优先优势
STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking
浏览论文内容
中文总结 AI 辅助
STAR-GRPO提出可靠性优先优势估计器,通过配对评估分离质量与学习影响,在组归一化前自调稳健拟合,有效抑制奖励黑客,提升规范质量并缩小代理-评判差异。
中文摘要 AI 辅助
当策略优化利用脆弱的奖励接口或过于宽松的代理目标时,会发生奖励黑客现象,即提高训练分数而不改善底层响应质量。这一现象在组相对策略优化中被放大:不受支持的奖励可能改变组基线并改变其他回放的更新,而事后或纯相对加权无法表示组级不确定性。我们提出自调锚定可靠性组相对策略优化(STAR-GRPO),一种基于同一回放配对评估的可靠性优先优势估计器。STAR将质量信号与其学习影响分离:分数分歧决定回放可靠性,相对可靠性在组归一化前进入自调稳健位置-尺度拟合,绝对组可靠性衰减所得的有界优势。分析建立了坐标和二阶矩界限,通过加权位置方程刻画精确居中,并给出对异常奖励的可靠性依赖衰减保证。我们在两个互补的奖励黑客机制中评估STAR-GRPO。在令牌接口利用中,STAR防止部署接口分数的失控优化,同时改善规范质量信号。在医学推理的规则代理过度优化中,STAR改善独立语义评估,缩小代理-评判者差异,并在优化同一任务代理时减少过度声称。综合这些结果表明,可靠性优先归一化为限制不受支持的奖励对组基线和策略更新的影响提供了一种原则性方法,同时保留任务奖励作为优化目标。
英文摘要
Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.
发表机构
- Peking University(北京大学)
- Beihang University(北京航空航天大学)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。