发表机构
University of California, Los Angeles; Amazon(加州大学洛杉矶分校; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对重复采样产生冗余尝试的问题,提出规划测试时扩展(PTTS),通过规划器生成大纲引导分支走向不同推理路径,以协调联合策略提升推理性能,在数学基准上显著提高pass@k。
AI 中文摘要
测试时扩展(Test-time scaling)通过并行分支被广泛用于提高具有挑战性的推理任务的性能。主流方法——重复采样(repeated sampling)——从单一策略中独立地抽取分支,这可能会产生冗余的尝试,从而限制了额外推理计算带来的收益。为解决这一局限,我们提出了规划测试时扩展(Planned Test-Time Scaling, PTTS),它用协调的联合策略取代独立采样:一个规划器(planner)为每个分支生成一个解决方案大纲,将分支引导至不同的推理路径,而一个执行器(executor)根据每个大纲生成完整的解决方案。形式上,我们证明PTTS严格推广了重复采样,并且在一种风格化的设定中,可证明地促进了对互补推理模式的覆盖,并产生更好的pass@k扩展。我们在强大的推理模型之上实例化PTTS,保持它们作为执行器固定,同时用PTTS推理替换重复采样以进一步增强测试时扩展。具体而言,我们开发了两种变体:PTTS-ZS提示模型在单次自回归传递中联合生成所有分支的大纲,而PTTS-RL直接针对pass@k奖励优化规划器,使用截断执行展开(truncated execution rollouts)以实现高效训练和更尖锐的奖励信号。在五个数学推理基准上,使用Qwen3-1.7B和4B,PTTS-ZS将pass@64相对于重复采样提高了最多6.7个百分点,而PTTS-RL进一步将增益提高到最多13.4个百分点。进一步分析表明,对不同推理路径的更广泛覆盖促成了这些增益。总体而言,PTTS通过协调推理分支提供了一个改进测试时扩展的通用框架,具有零样本和可训练的实例化,并产生了显著的性能提升。
英文摘要
Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce redundant attempts and thereby limit the gains from additional inference compute. To address this limitation, we propose Planned Test-Time Scaling (PTTS), which replaces independent sampling with a coordinated joint policy: a planner generates a solution outline for each branch, steering the branches toward distinct reasoning paths, and an executor produces a full solution conditioned on each outline. Formally, we show that PTTS strictly generalizes repeated sampling and, in a stylized setting, provably promotes coverage of complementary reasoning modes and yields better pass@k scaling. We instantiate PTTS on top of strong reasoning models, keeping them fixed as executors while replacing repeated sampling with PTTS inference to further enhance test-time scaling. Concretely, we develop two variants: PTTS-ZS prompts a model to jointly generate outlines for all branches in a single autoregressive pass, while PTTS-RL directly optimizes the planner against the pass@k reward using truncated execution rollouts for efficient training and a sharper reward signal. Across five mathematical reasoning benchmarks with Qwen3-1.7B and 4B, PTTS-ZS improves pass@64 over repeated sampling by up to 6.7 points, while PTTS-RL further increases the gain to up to 13.4 points. Further analysis indicates that broader coverage of distinct reasoning paths contributes to these gains. Overall, PTTS provides a general framework for improving test-time scaling by coordinating reasoning branches, with zero-shot and trainable instantiations that yield substantial performance gains.