arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25637cs.CL

AutoVerifier:面向基于参考的答案验证的残差引导非参数优化

AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification

Zebei Zhao, Zhihao Shi, Minqi Shi

首次发表
浏览论文内容

中文总结 AI 辅助

针对基于参考的答案验证中答案等价性依赖问题与评分标准的隐含偏置问题,提出残差引导的非参数优化方法AutoVerifier,从验证错误中学习偏置并经重放验证后升级规则,在四个基准上大幅超越现有最优验证器。

中文摘要 AI 辅助

基于参考的验证器对于评估推理模型、在具备可验证奖励的强化学习中提供准确的结果奖励至关重要。为提升验证准确率,已有研究探索了基于规则、基于模型和工具增强的验证器,用于检查不同答案形式之间的等价性。然而,诸如$1+3.14$与$1+π$这类答案形式的等价性可能取决于问题和评分标准。我们将这类隐含假设定义为验证器的归纳偏置。为解决这一挑战,我们提出AutoVerifier,这是一种残差引导的非参数优化方法,可从反复出现的验证器错误中学习这些偏置。具体而言,AutoVerifier将这些偏置记录在规则卡中,仅在重放验证未检测到直接回归时,才将其升级为代码模块或提示引导,确保被采纳的更新可审计、可编辑且可复用。在四个验证器基准上的实验表明,AutoVerifier大幅优于现有最优验证器。

英文摘要

Reference-based verifiers are important for evaluating reasoning models and providing accurate outcome rewards in reinforcement learning with verifiable rewards. To improve verification accuracy, prior work has explored rule-based, model-based, and tool-augmented verifiers for checking answer equivalence across diverse answer forms. However, the equivalence of answer forms such as $1+3.14$ and $1+π$ may depend on the question and scoring criterion. We frame such implicit assumptions as verifier inductive biases. To address this challenge, we propose AutoVerifier, a residual-guided non-parametric optimization method that learns these biases from recurring verifier errors. Specifically, AutoVerifier records these biases in rule cards and promotes them to code modules or prompt guidance only after replay validation detects no direct regressions, keeping accepted updates auditable, editable, and reusable. Experiments on four verifier benchmarks demonstrate that AutoVerifier outperforms state-of-the-art verifiers by a large margin.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑