arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

努力而非明智:推理模型无法在不同问题间合理分配测试时计算资源

Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions

Chenrui Fan, Yize Cheng, Ming Li, Yongyuan Liang, Tianyi Zhou, Soheil Feizi

arXiv 2608.07968首次发表:更新:

发表机构

University of Maryland; MBZUAI(马里兰大学; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出考试式评估框架发现推理模型无法在不同难度分值的问题间合理分配共享测试时计算预算,明确规划提示可使分配更均匀但仍不敏感于分值难度,全局预算分配是现有模型未掌握的独特能力。

AI 中文摘要

推理类语言模型越来越多地利用测试时计算来提升性能,但现有评估通常是针对单个问题的计算资源进行研究。然而当多个问题共享端到端的成本或延迟约束时,模型必须决定如何在这些问题间分配有限的推理计算资源。我们引入了一种考试式评估框架来研究这一情境,其中模型必须在不同难度和分值的问题间分配一个共享的 token 预算,以最大化其总分。在多个开源及前沿推理模型上,我们发现模型无法在不同难度和分值的问题间战略性地分配共享预算。模型的表现基本如同贪心序列求解器:它们按呈现顺序对问题排序,将计算资源前置分配给早期问题,且对分值不敏感,随着问题数量增加,这些倾向会愈发明显。明确的规划提示能让计算资源分配更均匀,但无法产生对分值或难度敏感的优先级排序。这一行为模式从数学推理延伸至代码推理。这些发现表明,全局预算分配是一种独特的能力,它未被传统的单问题评估所捕捉,仍是当前推理模型面临的一项挑战。

英文摘要

Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency constraint, models must decide how to divide limited inference compute among them. We introduce an exam-style evaluation framework for studying this setting, in which a model must distribute one shared token budget across questions with different difficulty and point values to maximize its total score. Across several open and frontier reasoning models, we find that models fail to allocate a shared budget strategically across questions of varying difficulties and values. Models behave largely as greedy sequential solvers: they prioritize questions by presentation order, front-load effort on early questions, and remain insensitive to value, with these tendencies becoming more pronounced as the number of questions grows. Explicit planning prompts spread compute more evenly but do not produce value- or difficulty-aware prioritization. The same behavioral pattern extends from mathematical to code reasoning. These findings establish global budget allocation as a distinct capability that is not captured by conventional per-question evaluation and remains a challenge for current reasoning models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑