arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07424cs.AI

CoBa:通过计算均衡路由实现高性价比的测试时扩展

CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing

Yan Zhou, Yue Ouyang, Kaiyang Zheng, Suncheng Xiang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出CoBa计算均衡路由策略,在固定推理预算下优化测试时计算分配,在多个数学推理基准上实现高准确率的同时大幅降低计算资源消耗。

中文摘要 AI 辅助

测试时扩展通常通过在单一维度上消耗更多计算资源来实现:采样更多解决方案、扩展思维链或应用更强的评估器。在固定推理预算下,这些选择存在竞争关系。本文将测试时推理建模为计算分配问题,系统需决定下一个计算单元应分配给生成、验证还是停止。我们提出CoBa(Compute-Balanced Routing,计算均衡路由策略),该策略首先获取一小部分候选,广泛应用低成本验证,再将不确定或高价值候选路由至更强的验证。在涵盖MATH-500、AIME 2024/2025、AMC 2023及过程符号推理的3129个示例生成器评估中,CoBa-Routed-Strong达到85.13%的宏观准确率,与自评估加权投票代理的85.20%在统计上匹配,同时使用的参数加权标记减少49.1%;其与best-of-16多数投票的宏观准确率差距在0.01个点以内,同时使用的参数加权标记减少58.9%,配对测试显示在更高成本下仍保留small best-of-16的优势。配对自举测试表明,相比单样本解码有显著提升,与池预言机的剩余差距则为更精准的路由提供了提升空间。对于局部推理系统,测试时扩展的核心问题转变为下一次计算应在何处发挥最大价值。

英文摘要

Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or applying a stronger evaluator. Under a fixed inference budget, these choices compete. This paper formulates test-time reasoning as a compute-allocation problem in which a system must decide whether the next unit of compute should be spent on generation, verification, or stopping. We introduce CoBa, a compute-balanced routing policy that first obtains a small set of candidates, applies cheap verification broadly, and routes uncertain or high-value candidates to stronger verification. On 3,129 example-generator evaluations spanning MATH-500, AIME 2024/2025, AMC 2023, and procedural symbolic reasoning, CoBa-Routed-Strong reaches 85.13% macro accuracy, statistically matching a self-evaluation weighted-voting proxy at 85.20% while using 49.1% fewer parameter-weighted tokens. It also matches best-of-16 majority voting within 0.01 macro-accuracy points while using 58.9% fewer parameter-weighted tokens; paired tests retain a small best-of-16 edge at substantially higher cost. Paired bootstrap tests show significant gains over single-sample decoding, while the remaining gap to the pool oracle exposes headroom for sharper routing. For local reasoning systems, test-time scaling becomes a question of where the next computation is most valuable.

补充信息

↑