发表机构
University of Chinese Academy of Sciences; Institute of Information Engineering, Chinese Academy of Sciences(中国科学院大学; 中国科学院信息工程研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ROUTEAUDIT通过契约条件识别和归因证书,实现预算多验证器路由的交互感知比较,在保留数据上验证了归因准确性和策略差异。
AI 中文摘要
自适应多验证器系统通常通过端点质量-成本差距进行比较,即使验证器目录、可用性、核算、信息过滤或评分器随策略变化也是如此。我们将验证器路由形式化为一个契约条件下的识别问题。该契约记录请求支持、验证器目录、实际可用性、资源核算、在线过滤和事后轨迹评分;匹配的路由对比仅改变策略坐标。ROUTEAUDIT在此契约中增加了三个可测量对象。契约格对每个允许的桥接顺序的坐标增量进行平均,并报告由此产生的归因及其路径敏感性。策略无关的响应磁带在自适应策略揭示不同观测时识别成对的顺序对比。对于不完全匹配,请求级边界使用任一仍可观测的潜在结果,并给出尖锐的有限总体区间。该协议在预言机加入之前提交已付费观测和账本事件,并为每次比较返回归因证书。在两个保留的原始尾部缓存上,匹配的静态SF+SA等于级联,将相对于完全静态的0.1797和0.1250的表观增益归因于验证器集优势。在1,319个保留的任务请求上,学习和RLVR研究报告的质量分别为0.9522和0.9553,而匹配静态为0.9484;RLVR-静态配对差异为+0.0068,请求配对区间为[0.0015,0.0122],训练种子按请求的层次区间为[0.0006,0.0131]。受控归因恢复产生路由平均绝对误差0.0011和端点重建误差0.0004。因子、桥接顺序和随机提供者研究评估了证书接口;RLVR在相同识别契约下提供了学习策略的压力测试。
英文摘要
Adaptive multi-verifier systems are commonly compared through endpoint quality-cost gaps, even when the verifier catalog, availability, accounting, information filtration, or scorer changes with the policy. We formulate verifier routing as a contract-conditioned identification problem. The contract records request support, verifier catalog, realized availability, resource accounting, online filtration, and post-trace scoring; a matched route contrast changes only the policy coordinate. ROUTEAUDIT adds three measurable objects to this contract. A contract lattice averages coordinate increments over every admissible bridge order and reports the resulting attribution together with its path sensitivity. A policy-independent response tape identifies paired sequential contrasts when adaptive policies reveal different observations. For incomplete matching, request-level bounds use whichever potential outcome remains observed and give a sharp finite-population interval. The protocol commits paid observations and ledger events before the oracle join and returns an attribution certificate for each comparison. On two held-out raw-tail caches, matched static SF+SA equals the cascade, assigning the apparent gains of 0.1797 and 0.1250 over full static to the verifier-set edge. On 1,319 held-out task requests, the learned and RLVR studies report quality 0.9522 and 0.9553 versus 0.9484 for matched static; the RLVR-static paired difference is +0.0068 with a request-paired interval $[0.0015,0.0122]$ and a training-seed-by-request hierarchical interval $[0.0006,0.0131]$. Controlled attribution recovery yields route mean absolute error 0.0011 and endpoint reconstruction error 0.0004. Factorial, bridge-order, and stochastic-provider studies evaluate the certificate interface; RLVR supplies a learned-policy stress test under the same identification contract.
Comments45 pages, 15 figures