长时程智能体任务中的一致规划-执行
Consistent Plan-Act for Long-Horizon Agentic Tasks
另 1 家 · 查看机构详情
- Nanjing University(南京大学)
- National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术全国重点实验室)
- Meituan(美团)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对长时程智能体任务中规划者与执行者状态不一致导致的协调失败,提出ConPAct方法,通过反馈矛盾并微调提升协调与性能,如MiniGrid成功率从38.6%升至54.4%。
中文摘要 AI 辅助
长时程智能体任务要求在动态环境的连续交互中具备强大的推理能力和高效执行能力。一种常见的方法通过分离的规划者和执行者角色,将高层规划与低层执行解耦。为了研究这些任务中的协调失败问题,我们提示两个智能体进行结构化状态断言,并通过程序化方式比较它们的报告以检测明确的矛盾。我们的分析揭示了关于同一任务相关状态事实的系统性分歧,我们将这一现象称为规划者-执行者状态不匹配。我们进一步发现,向智能体提供任务相关的状态信息可以减少不匹配,并改善协调和任务性能。基于对状态不匹配的系统性分析,我们提出了一致规划-执行(ConPAct),该方法将检测到的矛盾反馈给两个智能体,以形成一致的状态解释,并在精选的一致交互上对它们进行微调以实现更好的协调。ConPAct在多种环境和模型配置下提升了性能,例如,在MiniGrid中,使用GPT-5.6-sol/terra分别作为规划者和执行者时,成功率从38.6%提升至54.4%,这表明状态一致性可以指导推理时的修正和协调训练。
英文摘要
Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles. To investigate coordination failures in these tasks, we prompt both agents for structured state assertions and compare their reports programmatically to detect explicit contradictions. Our analyses reveal systematic disagreement about the same task-relevant state facts, a phenomenon we term planner-actor state mismatch. We further find that providing agents with task-relevant state information reduces mismatch and improves coordination and task performance. Based on the systematic analysis of the state mismatch, we propose Consistent Plan-Act (ConPAct), which feeds detected contradictions back to both agents to form consistent state interpretations and fine-tunes them on curated consistent interactions for better coordination. ConPAct improves performance across various environments and model configurations, e.g., increasing MiniGrid success rate from 38.6% to 54.4% with GPT-5.6-sol/terra as planner and actor respectively, demonstrating that state consistency can guide both inference-time correction and coordination training.