审计优先VAPO:不完美验证下的风险认证选择性更新
Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification
- University of Chinese Academy of Sciences(中国科学院大学)
- Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
审计优先VAPO通过分离方向准入与幅度控制,在不完美验证下认证选择性更新,以有限预算实现风险可控,并在基准上优于基线方法。
AI中文摘要:
不完美的验证器即使在裁剪和正则化约束其幅度的情况下,也可能分配有害的更新方向。我们引入了审计优先VAPO,它将离散的方向性准入与连续的幅度控制分离。一个仅观察的接受-申诉-弃权(不执行)策略使用有限的二次验证预算;其动作轨迹在干净标签加入之前被冻结。同时的有限样本界随后在预先声明的策略族上认证选定的有害风险、覆盖率和验证器调用率。条件Hoeffding-Azuma界考虑了共享预算引起的依赖性,而回滚或验证器更改会启动新的认证阶段。在准入之后,一个有界的信任-裁剪-KL执行器控制幅度。我们在两个推理基准上评估了两个模型,与静态RLVR、匹配随机选择、置信度阈值、噪声校正和验证器增强进行比较。在Qwen3.5-0.8B和GSM8K上,目标风险ρ=0.08时,RC-VAPO实现了74.1%的准确率,选定的有害风险为0.0697,覆盖率为0.4125,相对验证器成本为1.16倍。在匹配的覆盖率和更新幅度下,其与匹配随机的选定风险差异为-0.0260,配对95%区间为[-0.0364,-0.0157]。在不对称、置信度依赖和相关的验证器噪声下,证书在60次独立运行中的57次得到满足。这些比较将信息性方向选择与提议抑制、更新收缩和额外验证器计算隔离开来。
英文摘要:
Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite secondary-verification budget; its action trace is frozen before clean labels are joined. Simultaneous finite-sample bounds then certify selected harmful risk, coverage, and verifier-call rate over a predeclared policy family. Conditional Hoeffding-Azuma bounds account for the dependence induced by shared budgets, and rollout or verifier changes initiate a new certification stage. After admission, a bounded trust-clip-KL actuator controls magnitude. We evaluate two models on two reasoning benchmarks against static RLVR, matched-random selection, confidence thresholding, noise correction, and verifier augmentation. On Qwen3.5-0.8B and GSM8K at target risk $ρ=0.08$, RC-VAPO achieves 74.1% accuracy, selected harmful risk 0.0697, coverage 0.4125, and relative verifier cost $1.16\times$. At matched coverage and update magnitude, its selected-risk difference from matched random is -0.0260 with paired 95% interval $[-0.0364,-0.0157]$. Across asymmetric, confidence-dependent, and correlated-verifier noise, the certificate is satisfied on 57 of 60 independent runs. These comparisons isolate informative directional selection from proposal suppression, update shrinkage, and additional verifier computation.