验证器失效之处:RLVR中奖励信号的类别级审计
Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR
浏览论文内容
中文总结 AI 辅助
本文对RLVR的四个验证器开展类别级审计,发现验证器自验证率差异大,错误集中在空格标点,且不同验证器失效原因相反,数值级联验证器存在阶跃式接受差1个的错误答案的问题。
中文摘要 AI 辅助
带可验证奖励的强化学习(RLVR)与标准基准评估均依赖自动验证器,将自由文本答案转换为二元奖励。现有研究报告称,某评估工具仅接受自身真实答案的约94%,将此归咎于LaTeX解析,这是一个整体统计,未说明哪些答案形式消耗了错误预算。本文对此进行分解,采用变异测试针对验证器而非模型,生成经认证的等价答案变体,即通过构造保留数学含义的改写,因此任何拒绝都是可证明的假阴性,无需人工裁决。随后,在307420个裁决样本上,对四个广泛使用的验证器,按答案类别测量拒绝率,得出三点发现:1. 相同输入下的自验证率在53.8%至95.2%之间,差值达41.3个百分点,已发表数据仅描述一种实现而非任务,同一库的两种配置对49.9%的配对意见不一致;2. 剩余错误并非分布在解析类别中,而是集中在空格和标点,默认LaTeX配置下,这两类占符合合同失败的93.0%,尾点或换行是错误预算的主要来源;3. 将拒绝与执行失败分离后发现,聚合错误相似的验证器失效原因相反,参考数值级联因相对容差与尺度无关,以幅度阶跃函数形式接受“差1个”的错误答案,即10^4以下为0%,10^4及以上为100%。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) and standard benchmark evaluation both rely on an automatic verifier that turns a free text answer into a binary reward. Prior work reports that one evaluation harness accepts only about 94% of its own ground truth answers, blaming LaTeX parsing. That is an aggregate: it does not say which answer forms consume the error budget. We supply the decomposition. We apply metamorphic testing to the verifier rather than the model, generating certified equivalent answer variants, that is, rewrites that preserve mathematical meaning by construction, so that any rejection is a provable false negative needing no human adjudication. We then measure rejection per answer category across four widely used verifiers over 307,420 verdicts. We find three things. (1) Self validation ranges from 53.8% to 95.2% on identical inputs, a spread of 41.3 points. The published figure describes one implementation, not the task; two configurations of the same library disagree on 49.9% of pairs. (2) The residual is not spread across parsing categories but concentrated in whitespace and punctuation, which account for 93.0% of in contract failures for the default LaTeX configuration. A trailing period or newline dominates the budget. (3) Separating rejection from execution failure shows that verifiers with similar aggregate error fail for opposite reasons, and that a reference numeric cascade accepts off by one wrong answers as a step function of magnitude, from 0% below 10^4 to 100% at or above, because its relative tolerance is scale invariant.