arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重采样还是重新路由?大语言模型的预算感知测试时模型选择

Resample or Reroute? Recoverable Stopping Debt Without Identified Action Selection

Teng-Ruei Chen

arXiv 2607.08665首次发表:更新:

发表机构

Institute of Bioinformatics and Systems Biology, National Yang Ming Chiao Tung University; Krixvon(国立阳明交通大学生物信息与系统生物学研究所; 克瑞克斯冯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型测试时模型选择问题,提出预算感知测试时模型选择方法,通过在重采样和重新路由间分配预算单位最大化预期正确性,提出RoR分配策略,实验表明该策略在成本-质量方面优于多种基线。

AI 中文摘要

大语言模型(LLMs)之间的路由在响应质量和服务成本之间进行权衡,这是由于已部署路由器与每个实例的理想模型之间存在差距。近期分析表明,测试时重采样可恢复单个提交路由器无法捕获的每个实例选择空间,但该保证仅在配备正确性标签和无约束预算的理想化模型下成立,而实际部署系统并不具备这些条件。此前没有工作将重采样已提交模型和重新路由到替代模型视为单个查询成本预算的竞争用途。因此,本文提出预算感知测试时模型选择,即在给定每个查询预算和不完美验证器的情况下,在重采样和重新路由之间分配预算单位,以最大化预期正确性。提出了一种由每单位成本的估计边际正确性驱动的在线重采样或重新路由(RoR)分配策略,其行为基于选择和采样之间的可恢复性不对称。在四个不同难度基准上对来自11个模型的开放权重池新生成的多抽头正确性张量进行回放实验表明,相对于单路由、单提交路由器、预算感知最佳K、级联和随机分配基线,所提出的RoR策略在测试池中实现了良好的成本-质量帕累托前沿,在最异构的基准上增益最大;消融实验进一步表明,增益是由验证器控制的,随着验证器质量下降而缩小,并且在提供商价格向量和无标签协议验证器下的稳健性回放描绘了结论适用的范围。

英文摘要

After a weak verifier accepts a large-language-model response, a second call may resample or reroute. Because correctness is hidden, action selection is an identification problem. We order three gates: recoverable stopping debt, two-sided FIT action support, and held-out value from an outcome-blind selector. In a pinned 152-query MBPP+ experiment, a Qwen2.5-14B Base-only false-positive stop leaves +2.592 percentage points of Qwen2.5-7B recovery (query-cluster 95% interval [+1.618, +3.664]). Separately, after 7B Base-test rejection, fixed escalation to 14B exceeds leave-one-out 7B resampling by +2.882 points [+0.931, +5.201]; this is fixed-action ranking, not conditional selection. An all-episode audit produces a +2.697-point realized-maximum gap, but for two actions this statistic equals (1/2)E|Delta| - (1/2)|E Delta| and contains no observable-history term. It lies inside an exact-fold exchangeable reference (mean +3.158; 95% interval [+2.434, +3.947]). The audit unconditionally acts on 1,520 episodes: 1,240 observable stops and 280 verifier rejections; 198 stops are evaluator-only false positives. Neither tested outcome-blind controller improves on fixed rerouting. A separate LiveCodeBench ladder has all-zero FIT action advantages despite exclusive TEST rescues. A preregistered BigCodeBench support gate then finds only 23/19 and 22/19 signed episodes/queries against minima of 25/20, so L1-L4, DEV, and TEST stay unopened. Stopping debt exists, but current evidence does not identify when to resample rather than reroute.

Comments33 pages (IEEEtran two-column) incl. 14-page supplement; 14 figures, 15 tables. v3: substantive methodological reconstruction of v1-v2. The budget-allocation, learned-router, cost-oracle and broad Pareto claims are not retained; replaced by a three-gate identification framework with fresh MBPP+, LiveCodeBench and preregistered BigCodeBench evidence. v2: corrected Phi-4 cost (14.7B)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑