DeepPlanning: 用可验证约束基准测试长周期代理规划
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DeepPlanning是一个针对长周期代理规划的基准测试,通过多日旅行和多产品购物任务,评估代理在全局约束优化和局部约束推理中的能力。
AI中文摘要:
尽管代理评估已转向长周期任务,但大多数基准仍强调局部、步骤级推理,而非要求真实规划能力的全局约束优化(例如时间与财务预算)。同时,现有LLM规划基准未能充分代表现实世界中典型的主动信息收集和细粒度局部约束。为此,我们引入DeepPlanning,一个针对实际长周期代理规划的挑战性基准。它包含多日旅行规划和多产品购物任务,要求主动获取信息、局部约束推理和全局约束优化。对DeepPlanning的评估显示,即使是前沿的LLM代理也难以解决这些问题,突显了可靠显式推理模式和并行工具使用在实现更好效果-效率权衡中的重要性。错误分析进一步指出了改进长周期规划LLM的有希望方向。我们开源代码和数据以支持未来研究。
英文摘要:
While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands genuine planning ability. Meanwhile, existing LLM planning benchmarks underrepresent the active information gathering and fine-grained local constraints typical of real-world settings. To address this, we introduce DeepPlanning, a challenging benchmark for practical long-horizon agent planning. It features multi-day travel planning and multi-product shopping tasks that require proactive information acquisition, local constrained reasoning, and global constrained optimization. Evaluations on DeepPlanning show that even frontier agentic LLMs struggle with these problems, highlighting the importance of reliable explicit reasoning patterns and parallel tool use for achieving better effectiveness-efficiency trade-offs. Error analysis further points to promising directions for improving agentic LLMs over long planning horizons. We open-source the code and data to support future research.