arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeReAct:面向可靠AI智能体的分解推理与行动

DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents

Ajay Vohra, Tao Chen, Neeti Narayan, Caron Zhang

arXiv 2610.02351首次发表:更新:

发表机构

Amazon; Apple(亚马逊; 苹果)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DeReAct通过外部化Critic和Context Manager两个门控策略,分解推理与行动,提升较弱AI智能体的可靠性,在GAIA和SWE-bench上显著提高Pass@1,同时保留接地性优势。

AI 中文摘要

基于ReAct的智能体通常依赖单一的LLM策略来提出行动、与环境交互并决定任务何时完成。这种耦合使得行动授权和完成控制难以独立执行,导致错误得以传播,且未经支持的任务完成声明会终止执行。我们提出了DeReAct,一种模块化智能体架构,它外部化了两个门控策略:一个Critic在执行前验证提议的行动,以及一个Context Manager重建环境支持的\ extsc{State}并认证任务完成。在GAIA和SWE-bench Verified上,DeReAct对较弱的Brain模型在Pass@1上的提升最大,其中Qwen3-Coder-480B提升了6.5至7.0个百分点,Claude Sonnet 4.5提升了4.2至5.2个百分点;随着Brain能力的增强,提升幅度逐渐减小。轨迹和消融分析表明,当目标失败足够普遍且门控策略本身足够有效时,外部门控是有效的。使用Claude Opus 4.5时,Pass@1与ReAct相当,而DeReAct产生了更完整证据和更符合约束的轨迹,这表明完成控制可以用更早的终止来换取更强的接地性。总体而言,DeReAct改进了较弱的智能体,同时随着模型增强保留了接地性优势。

英文摘要

ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action authorization and completion control difficult to enforce independently, allowing errors to propagate and unsupported completion claims to terminate execution. We introduce DeReAct, a modular agent architecture that externalizes two gating policies: a Critic that validates proposed actions before execution, and a Context Manager that reconstructs an environment-supported \textsc{State} and certifies task completion. Across GAIA and SWE-bench Verified, DeReAct improves Pass@1 most for weaker Brain models, with gains of 6.5--7.0 points for Qwen3-Coder-480B and 4.2--5.2 points for Claude Sonnet~4.5; gains diminish as Brain capability increases. Trajectory and ablation analyses show that external gating is effective when targeted failures are sufficiently prevalent and the gating policy is itself sufficient. With Claude Opus~4.5, Pass@1 remains comparable to ReAct, while DeReAct produces more evidence-complete and constraint-satisfying trajectories, indicating that completion control can trade earlier termination for stronger grounding. Overall, DeReAct improves weaker agents while retaining grounding benefits as models strengthen.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑