发表机构
IU Bloomington(印第安纳大学布卢明顿分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出发现-执行框架,通过短预算运行预测数学推理随计算扩展的曲线,并在奥林匹克问题上验证,表明测量条件执行能提供短预算成功率无法捕捉的更长推理信息。
AI 中文摘要
测试时计算可以改善数学推理,但短预算运行能否预测数学推理如何随额外计算扩展?我们引入了一个发现-执行(DE)框架,该框架通过策略发现和条件执行的卷积来预测聚合的留出扩展曲线。从独立的短预算尝试和基于oracle草稿条件的运行中,该框架估计在替代计算分配下沿留出推理轨迹的累积成功。我们在35道新奥林匹克问题以及IMO-ProofBench Advanced中的非几何问题上评估了四个模型。在DE框架下,接近饱和的执行预测几何扩展,正如GPT模型所观察到的那样。对于Claude Opus 4.8,结合实测执行显著改善了在单臂和双臂分配下对留出预测的几何外推。作为次要应用,正则化DE(R-DE)决定继续或重启产生的平均遗憾低于最佳模型特定回顾性策略。这些结果表明,测量条件执行提供了关于更长推理的信息,而短预算成功率并不总能捕捉到这些信息。
英文摘要
Test-time compute can improve mathematical reasoning, but can short-budget runs predict how mathematical reasoning scales with additional compute? We introduce a Discovery--Execution (DE) framework that predicts the aggregate held-out scaling curves through a convolution of strategy discovery and conditional execution. From independent short-budget attempts and oracle-sketch-conditioned runs, the framework estimates cumulative success along held-out reasoning trajectories under alternate compute allocations. We evaluate four models on 35 fresh Olympiad problems and non-geometry problems from IMO-ProofBench Advanced. Under the DE framework, near-saturated execution predicts geometric scaling, as observed for the GPT models. For Claude Opus 4.8, incorporating measured execution substantially improves held-out forecasts over geometric extrapolation across one- and two-arm allocations. As a secondary application, regularized DE (R-DE) decisions to continue or restart yield lower average regret than the best model-specific retrospective policy. Together, these results show that measuring conditional execution provides information about longer reasoning that short-budget success rates do not always capture.
CommentsAccepted in NeuRIPS MATH-AI Workshop 2026