发表机构
Beijing Jiaotong University(北京交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究RLVR中奖励套件自然误报问题,通过预注册双臂因果对比实验,用GRPO在相同任务等条件下对比原始与强化测试奖励,发现平均留出效应有界,奖励误报与静态审计相关,还对评判者自身进行测试,表明廉价审计可定位暴露,强化奖励能消除测量膨胀。
AI 中文摘要
用作代码RLVR奖励的测试套件存在自然误报:即每个任务中持续存在的不对称错误,每次出现时都会接受相同的错误程序,这与现有抗噪声分析所假设的对称或重采样噪声不同。我们在一个已部署的套件上进行了预注册的双臂因果对比:在相同的MBPP任务、种子和计算资源上使用GRPO,分别由原始的MBPP测试(有漏洞)和MBPP+额外测试(强化)作为奖励。另外两个系列在数据存在之前预注册冻结的情况下重复了该设计。平均留出效应是有界的:在预注册的1.5分 margin 下非劣(差距0.20分,单侧95%上限0.75分)。奖励的误报质量与训练前计算的廉价静态漏洞审计相关(斯皮尔曼相关系数0.80),且注册的训练侧测试使漏洞层误报率比干净任务高43.8分。按照有签名的、人工裁决的规则审核每个奖励的误报发现,有大量经验证的真正错误代码残留:记录加权后为47.57%;两个重复系列也重现了很大比例。为真实错误支付奖励,而非仅仅为套件工件。机制证据与预先存在的错误模式选择而非学习利用一致:在我们的观察范围内误报发生率没有增加,未经训练的基础模型在有漏洞的过滤器下已经产生相同的错误输出。然后我们将同样的方法应用于前沿评判者自身:对于他们自己的误报,他们自我评估很弱,同一作者的测试未解决,甚至我们测试的得分最高的读者在较弱策略的错误上的得分也远低于其得分——关于MBPP的两个主题,对前沿模型一般没有任何许可。廉价的静态审计在训练前定位暴露;强化奖励消除了测量膨胀,尽管在此处它几乎没有带来能力提升。
英文摘要
The test suites used as RLVR rewards for code have natural false positives: per-task, persistent, asymmetric errors that accept the same wrong programs every time they appear, unlike the symmetric or resampled noise assumed by existing noise-robustness analyses. We run a preregistered two-arm causal contrast on a deployed suite: GRPO on identical MBPP tasks, seeds, and compute, rewarded by the original MBPP tests (leaky) versus the MBPP+ extra tests (hardened). Two further families replicate the design under a preregistration frozen before their data existed. [C] The average held-out effect is bounded: non-inferior under a preregistered 1.5-pt margin (gap 0.20 pt, one-sided 95% upper bound 0.75 pt). [C] Rewarded false-positive mass tracks a cheap static leakiness audit computed before training (Spearman 0.80), and the registered train-side test puts the leak-stratum FP share +43.8 pt above clean tasks. [E] Auditing every rewarded FP under signed, human-adjudicated rules finds a large residual of verified genuinely wrong code: 47.57% record-weighted; both replication families reproduce a large share. The reward paid for real bugs, not merely suite artifacts. [E] Mechanism evidence is consistent with selection of pre-existing error modes rather than learned exploitation: FP incidence does not grow within our horizon, and untrained base models already produce the same wrong outputs under the leaky filter. We then turn the same instrument on the frontier judges themselves: on their own false positives they self-assess only weakly, a same-author test is unresolved, and even the highest-scoring reader we probe stays far below its score on a weaker policy's errors -- two subjects on MBPP, licensing nothing about frontier models in general. A cheap static audit locates exposure before training; hardening the reward removes the measurement inflation, though here it buys little capability.
Comments37 pages, 2 figures. Code, frozen data, and complete audit records: https://github.com/toffee-desuwa/rlvr-leaky-suite