递归智能体推理
Recursive Agentic Reasoning
- Google(谷歌)
- University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文将测试时推理方法统一为递归算子,通过共享框架对比发现BRANCH算子在多场景下提升准确率,且配对评估是测试时计算评估的标准协议。
中文摘要 AI 辅助
测试时推理方法(如迭代精化、问题分解、重复采样)常被孤立评估,导致难以在不同模型、基准和评估流程中对比其增益。本文将这些方法统一视为智能体推理轨迹上的递归算子:GROW(加深单一推理路径)、PRUNE(分解并重构问题)、BRANCH(采样替代推理路径并从中选择)。在共享框架(提示词、token预算、评分代码完全一致)下,将这三种算子与单轮思维链基线对比,涉及5个基准、3个前沿模型,共14种模型-基准设置、49327个评分项、151876次模型调用。结果显示,BRANCH在全部14种设置中均提升准确率,平均增益5.98个百分点,在12种设置中表现最优;GROW平均增益2.18个百分点,在2种设置中降低性能;PRUNE平均提升0.94个百分点。分析表明,BRANCH的优势不仅源于探索多推理路径,还能从截断中恢复,其增益与基线的空输出、预算耗尽输出的比例强相关(相关系数r=0.72)。这些结果削弱了“不同问题需选择不同测试时推理算子”的假设,在该抽象层面,重复分支始终占优。最后,本文指出非配对评估及将评分流程失败视为模型错误,会显著改变甚至反转对比结论,推动配对评分成为测试时计算评估的标准协议。
英文摘要
Test-time reasoning methods such as iterative refinement, decomposition, and repeated sampling are often evaluated in isolation, making their gains difficult to compare across models, benchmarks, and evaluation pipelines. We introduce a unified view of these methods as recursion operators over an agent's reasoning trace: GROW, which deepens a single reasoning path; PRUNE, which decomposes and recomposes the problem; and BRANCH, which samples alternative reasoning paths and selects among them. We evaluate all three operators against a single-pass chain-of-thought baseline under a shared harness with identical prompts, token budgets, and grading code. Across five benchmarks and three frontier models, comprising 14 model-benchmark settings, 49,327 graded items, and 151,876 model calls, BRANCH improves accuracy in all 14 settings by an average of 5.98 percentage points and is the best-performing operator in 12. In contrast, GROW yields a mean gain of 2.18 points and degrades performance in two settings, while PRUNE improves accuracy by 0.94 points on average. Analysis shows that BRANCH's advantage arises not only from exploring multiple reasoning paths, but also from recovering from truncation: its gains strongly correlate with the baseline rate of empty, budget-exhausted outputs (r = 0.72). These results weaken the hypothesis that different problems require routing among test-time reasoning operators; at this level of abstraction, repeated branching is consistently dominant. Finally, we show that unpaired evaluation and treating scoring-pipeline failures as model errors can materially change, and even reverse, comparative conclusions, motivating paired scoring as a standard protocol for test-time-compute evaluation.