arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型能从测试时计算中获得多少收益?

How Much Can Language Models Gain from Test-Time Computation?

Bangji Yang, Jingyuan Li, Jiajun Fan, Yi Evie Zhang, Ruihan Guo, Hongbo Ma, Neil He, Chumeng Liang, Qinglong Zheng, Zhanghan Ni, Ge Liu

arXiv 2610.01110首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; University of Washington; Tsinghua University(伊利诺伊大学厄巴纳-香槟分校; 华盛顿大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SELF-POT基准,统一预算下评估测试时计算收益,发现选择规则和失败处理显著影响性能,公开示例选择提升编程正确率并节省成本。

AI 中文摘要

测试时计算能在多大程度上提升语言模型,其成本又是多少?测试时扩展被广泛提议作为更大模型的替代方案,但现有比较大多一次只评估一个领域,且很少将选择过程计入预算。我们引入了SELF-POT,一个基准测试和评估框架,用于衡量模型在竞赛数学、竞争性编程和智能体工作流中的测试时潜力。SELF-POT将静态任务中的候选覆盖与最终准确性分离,跟踪修订下的正确性转变,并在智能体环境中衡量协议完成情况以及任务成功。在统一预算规则下,它将直接推理与在直接预算的固定倍数下的并行采样和自我修订进行比较,并将每次模型调用(包括选择和评判)以美元计费。这种设计支持两种比较:模型从额外推理中获得的收益,以及一个成本更低的模型通过额外推理与更强模型的对比。在350个密封任务上对五个低成本推理模型进行测试,以Claude Opus 5.5直接推理作为参考,收益取决于领域、选择规则和失败处理方式。当我们重放保留的编程候选池时,公开示例选择将正确提交从500个计划单元中的376个提高到453个,同时在各模型上节省了12-49%的逻辑API成本;而仅仅在评判失败时保留可用候选,就能以不变的成本恢复61个提交。在相同的数学池上,带回退的评判产生186个正确提交,而投票为182个,同时投票节省了12-21%的逻辑API成本。这些受控重放展示了选择和失败处理如何改变从相同生成候选中所实现的收益,并量化了模型评判者的边际价值。

英文摘要

How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-POT separates candidate coverage from final accuracy on static tasks, tracks correctness transitions under revision, and measures protocol completion alongside task success in agentic environments. Under a unified budget rule, it compares Direct inference with parallel sampling and self-revision under fixed multiples of the Direct budget, and charges every model call, including selection and critique, in dollars. This design supports two kinds of comparison: the gain a model obtains from additional inference, and a lower-cost model with additional inference against a stronger model. Across five low-cost reasoning models on 350 sealed tasks, with Claude Opus 5.5 Direct as the reference, the returns depend on the domain, the selection rule, and failure handling. When we replay the retained programming candidate pools, public-example selection raises correct submissions from 376 to 453 of 500 scheduled cells while saving 12-49% of logical API cost across models, and simply retaining an available candidate when judging fails recovers 61 submissions at unchanged cost. On identical mathematics pools, judging with fallback yields 186 correct submissions versus 182 for voting, while voting saves 12-21% of logical API cost. These controlled replays show how selection and failure handling change the gains realized from the same generated candidates, and they quantify the marginal value of a model judge.

CommentsCorrected a typo in an author's name. No changes to the paper content

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑