GRPO 组内验证器错误是否独立?来自 Qwen2.5 生成结果的经验证据
Are Verifier Errors Independent Within a GRPO Group? Evidence from Qwen2.5 Rollouts
AI总结:
本研究通过 Qwen2.5 生成结果发现 RLVR 中同一验证器对同一提示的多个生成结果评分存在显著组内相关性,影响有效样本量,需考虑提示与答案形式的影响。
AI中文摘要:
基于可验证奖励的组强化学习(RLVR)使用自动验证器对每个提示的多个生成结果进行评分。基于独立验证器错误的分析可能忽略与共享答案格式相关的依赖性。我们在 MATH、GSM8K 和 DeepMath-103K 上由 Qwen2.5-1.5B 生成的 24,998 组每组八个生成结果中研究了这种依赖性。我们估计合并的组内验证器错误相关性为 0.530(95% 置信区间:0.500--0.560)。在可交换错误模型下,这对应于八生成结果组的设计效应调整后的有效样本量为 1.70。依赖性在不同答案形式间差异显著:分数、根式、符号表达式和区间比单位标注和百分号表现出更强的聚类。在四种基于规则的验证器配置中重放组相对优势,识别出最多 0.83% 的组存在至少一个优势符号不一致。因为一个组是对一个提示的重复采样,这种组内聚类可能反映共享的提示难度以及共享的答案形式,我们未尝试将两者分开。与跨多个评估者的相关判断研究不同,我们的分析检查了由同一验证器评分的生成结果之间的依赖性。这些发现促使对验证器噪声进行提示和答案形式感知的分析,而不是仅基于聚合错误率的表征。
英文摘要:
Group-based reinforcement learning with verifiable rewards (RLVR) scoresmultiple completions per prompt using automatic verifiers. Analysesbased on independent verifier errors may overlook dependence associatedwith shared answer formats. We investigate this dependence in24,998 groups of eight completions generated by Qwen2.5-1.5B onMATH, GSM8K, and DeepMath-103K. We estimate a pooled within-groupverifier-error correlation of 0.530 (95% confidence interval:0.500--0.560). Under an exchangeable-error model, this correspondsto a design-effect-adjusted effective sample size of 1.70 for aneight-completion group. Dependence varies substantially across answerforms: fractions, radicals, symbolic expressions, and intervals exhibitstronger clustering than unit annotations and percent signs. Replayinggroup-relative advantages across four rule-based verifier configurationsidentifies at least one advantage-sign disagreement in up to 0.83% ofgroups. Because a group is repeated sampling for one prompt, thiswithin-group clustering may reflect shared prompt difficulty as well asshared answer form, and we do not attempt to separate the two here.Unlike studies of correlated judgments across multiple evaluators, ouranalysis examines dependence across completions scored by the sameverifier. These findings motivate prompt- and answer-form-aware analysesof verifier noise rather than characterizations based solely onaggregate error rates.