arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.03436cs.LG

路由差距中有多少是真实的?将路由器到预言机的差距分解为可重现的专家优势和单次抽取标签噪声

What Does a Routing Oracle Measure Under Stochastic Decoding? Coupling, Scorer Choice, and Single-Commit Ceilings

Teng-Ruei Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型中路由差距,将预言机期望分解为可重现部分和选择下限,证明单次抽取标签噪声是差距一部分,提出多样本预言机评估协议,多数差距是可恢复专家优势。

中文摘要 AI 辅助

大语言模型间的路由有望以低成本实现更好质量,受限于学习到的路由器与每个实例预言机之间的差距。但预言机由单个正确性标签计算得出,具有随机性。我们从结构上重新审视该问题,将预期的每个实例预言机分解为可重现的单提交余量和非负单提交选择下限。主要结果是恢复不对称性,测试时采样可恢复下限,而单个提交路由器无法闭合。我们发布了用于路由基准测试的多样本预言机评估协议。

英文摘要

Routing benchmarks often compare a policy that commits to one model before seeing its response with a hindsight oracle that credits any correct recorded output. Under stochastic decoding these are different decision classes. For success marginals $p_{im}$ under a frozen query--model--decoder--scorer protocol, we distinguish the clairvoyant single-commit ceiling $R_i=\max_m p_{im}$, product-coupling union $U_i^\perp=1-\prod_m(1-p_{im})$, and premium $Δ_i^\perp=U_i^\perp-R_i$. The marginals alone identify exactly the union interval $[R_i,\min\{1,\sum_m p_{im}\}]$: zero premium is attainable over compatible couplings, not established for an actual deployment. If generation is independent across models, product is the specified protocol's union probability. We audit frozen correctness tensors from 11 open models and 30 archived responses per query--model cell on GSM8K, MATH-500, and GPQA-Diamond. Full-pool display-channel product premiums are 0.371, 3.542, and 5.100 percentage points; the premium intervals from empirical marginals are $[0,0.473]$, $[0,4.787]$, and $[0,7.744]$. Their zero lower endpoints are algebraic. All retained scorers, eight finite-draw paths, and all 2,047 nonempty subpools expose scorer, estimator, and pool sensitivity, not confidence intervals. The GSM8K and GPQA display scorers were developed after limited output inspection. A separate retrospective held-out policy illustration instantiates the policy-specific gap decomposition without establishing new-data generalization. A limited reference-based human check supports scorer agreement only on definite-consensus subsets. Hash-bound evidence supports number checks, not end-to-end reproduction. The contribution is a measurement contract for interpreting oracle gaps conditional on coupling, decision class, scorer, pool, and finite draws, not a population effect or an equal-cost routing gain.

发表机构

  • Institute of Bioinformatics and Systems Biology, National Yang Ming Chiao Tung University(国立阳明交通大学生物信息与系统生物学研究所)
  • Krixvon AI(Krixvon人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑