arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15303cs.AI

发散-收敛推理:通过结构化解决方案合成扩展测试时计算

Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis

Bo Wen, Yuhao Chen, Erhan Bilal, Carla Agurto Rios, Chen Wang, Junchen Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出发散-收敛推理(DCR),引入递归DCR方法,利用分歧信息优化测试时计算分配,在AIME 2024、2025数据集上实现高准确率且减少计算量,揭示LLM测试时推理的扩展规律。

中文摘要 AI 辅助

测试时计算可大幅提升大型语言模型(LLM)的推理性能,但额外计算的作用方式与时机仍未被充分理解。本文研究发散-收敛推理(DCR),这是一种简单的两阶段原语,包括生成多个候选解决方案的探索阶段,以及后续的收敛协调阶段。我们提出三项核心结果:第一,即使仅一次协调步骤也能可靠放大正确的少数观点:在多个数据集上,当探索阶段输出的正确答案占少数(多数投票法失效的场景)时,DCR常能恢复正确答案。第二,我们引入递归DCR,这是一种迭代分析分歧并分配额外测试时计算的自回归协调系统。递归DCR达到比固定计算基线更高的准确率——在AIME 2024上达93.3%,在AIME 2025上达92.0%,同时平均计算量减少约27%,证明针对性资源分配优于均匀扩展。第三,我们通过一个简单的无训练分散度指标分析探索输出间的分歧,该指标揭示了分歧与测试时收益的结构化关系:在DCR有效的场景中,探索输出间的更高分歧与协调带来的更大准确率提升相关。综上,这些结果表明,常被视为噪声的分歧可被系统地利用以改进测试时推理,并揭示了智能体式LLM系统的新兴扩展规律。

英文摘要

Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood. We study Divergent-Convergent Reasoning (DCR), a simple two-phase primitive consisting of an exploration phase that generates multiple candidate solutions followed by a convergent reconciliation phase. We present three core results. First, we show that even a single reconciliation step can reliably amplify correct minority reports: across datasets, DCR often recovers the correct answer when correct exploration outputs are in the minority, a regime where majority voting fails. Second, we introduce recursive DCR, an autoregressive reconciliation system that iteratively analyzes disagreements and allocates additional test-time compute. Recursive DCR achieves higher accuracy than fixed-compute baselines-reaching 93.3% on AIME 2024 and 92.0% on AIME 2025-while using roughly 27% less compute on average, demonstrating that attentive resource allocation is superior to uniform scaling. Third, we analyze disagreement among exploration outputs via a simple, training-free dispersion metric. Dispersion reveals a structured relationship between disagreement and test-time gains: in regimes where DCR is effective, higher disagreement among exploration outputs is associated with larger accuracy improvements from reconciliation. Together, these results show that disagreement, often viewed as noise, can be systematically exploited to improve test-time reasoning and reveal emerging scaling laws for agentic LLM systems.

发表机构

  • School of Computing, Queen’s University(女王大学计算学院)
  • IBM T.J. Watson Research Center(IBM T.J.沃森研究中心)
  • University of Chicago(芝加哥大学)
  • Tensormesh Inc.(Tensormesh公司)

机构由 AI 辅助整理,请以论文原文为准。

↑