arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ParaTempo:基于时间置信度的高效并行推理

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

Xuteng Zhang, Wenhao Zeng, Xiaodong Gu, Chao Hu, Haotian Lin, Yuling Shi, Min Wang, Beijun Shen

arXiv 2608.16425首次发表:更新:

AI 中文总结

ParaTempo 是一种基于时间置信度的无训练异步并行推理框架,可自适应分配计算资源,在保持推理准确率的同时降低延迟与 token 用量,且时间置信度性能优于其他信号。

AI 中文摘要

并行推理通过探索多条解路径提升大型推理模型的准确率与鲁棒性,但其计算成本随推理深度和分支数量增长。现有管理并行路径的方法通常依赖最终答案共识、局部 token 置信度或孤立的中间探测,但这些信号往往存在延迟、与实际推理进展关联弱,或噪声过大,难以用于动态的分支级控制。为解决这些局限,我们提出 ParaTempo,一个无需训练的异步并行推理框架。ParaTempo 由时间置信度驱动,这是一种针对答案空间收敛性的分支局部度量。每个分支会被定期探测以获取暂定答案概率分布,时间置信度量化近期中间探测对主导答案的集中程度。一旦积累了足够证据,ParaTempo 便基于这单一信号驱动全部控制流程:剪枝低置信度分支、提前退役持续承诺主导答案的分支、通过分叉新分支释放被重新分配的计算资源,当置信度加权投票集中时全局停止生成。无需在推理轨迹间同步,ParaTempo 基于分支级收敛性自适应分配计算资源。在具有挑战性的数学和科学推理基准上的实验显示,ParaTempo 将平均延迟降低 21.8%-32.2%,总 token 用量降低 18.1%-30.3%,同时保持具有竞争力的准确率。此外,时间置信度相比 token 级和瞬时信号,表现出更强的时间稳定性和对未来分支收敛的预测能力。

英文摘要

Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managing these parallel paths typically rely on final-answer consensus, local token confidence, or isolated intermediate probes. However, these signals are often delayed, weakly tied to actual reasoning progress, or too noisy for dynamic, branch-level control. To address these limitations, we introduce ParaTempo, a training-free asynchronous parallel reasoning framework. ParaTempo is driven by temporal confidence, a branch-local measure of answer-space convergence. Each branch is periodically probed for a tentative answer probability distribution, and temporal confidence quantifies how sharply the recent intermediate probes concentrate on a dominant answer. Once sufficient evidence has accumulated, ParaTempo drives its entire control process from this single signal: low-confidence branches are pruned, branches that persistently commit to their dominant answer are retired early, freed computation is reallocated by forking new branches, and generation stops globally once the confidence-weighted vote concentrates. Without requiring synchronization among reasoning trajectories, ParaTempo adaptively allocates computation based on branch-level convergence. Experiments on challenging mathematical and scientific reasoning benchmarks show that ParaTempo reduces average latency by 21.8-32.2% and total token usage by 18.1-30.3% while maintaining competitive accuracy. Moreover, temporal confidence exhibits stronger temporal stability and predictive power for future branch convergence than token-level and instantaneous signals.

CommentsCode and dataset are available at https://github.com/ScottZhang812/ParaTempo

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑