发表机构
Meta Superintelligence Labs(Meta超级智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出智能体元推理框架,通过控制器显式管理执行决策,在长时程任务中显著提升性能,如ProgramBench上GPT-5.5达71.5%,并随预算增长持续改进。
AI 中文摘要
随着智能体处理更长、更复杂的问题,控制执行过程本身成为一项任务。运行中的每一步都会带来新的控制选择,例如基于哪些部分工作继续推进、是否重新开始,或何时停止。我们引入了智能体元推理,这是一种推理时框架,将这些选择转化为显式且结构化的推理过程。工作线程负责任务级计算,而控制器则整合运行中已建立的成果,探索下一步选项,评估在剩余预算下每个选项的价值,并从持久记忆中提取上下文来分派所选工作。在决策之间,控制器仅携带运行的紧凑记录,而非重放完整历史。我们的基线涵盖生产级编码智能体和研究框架,以及使用相同工作线程和计算预算额度的直接控制智能体。在ProgramBench(通过程序重建测试长时程智能体能力)上,元推理在GPT-5.5下达到71.5%,而Codex为58.0%;在Opus 4.8下达到67.2%,而Claude Code为65.5%。在其他基准上,涵盖抽象推理、多领域长时程推理和证明生成,元推理相比直接控制平均提升3.6至4.2个百分点,跨三个前沿模型。在直接控制趋于平稳的测试预算范围内,元推理持续改进,尽管其开销在小预算下可能造成损害。工件图分析显示,元推理更多复用早期工作,在大多数设置下对正确解的覆盖率更高,并在最终选择上呈现非均匀增益。这些结果表明,随着智能体扩展到更长的运行,将计算投入结构化控制变得更为重要。
英文摘要
As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.