发表机构
Wuhan University; Hong Kong University of Science and Technology(武汉大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出LifePlanner基准测试,结合社交媒体数据评估LLM智能体的地理空间规划能力,发现前沿LLM在简单检索中表现好但复杂规划通过率仅40.2%,失败源于证据获取、工具使用及约束整合问题,需改进实际规划能力而非单纯扩大模型规模。
AI 中文摘要
地理空间规划(如行程设计)是LLM智能体的现实测试平台,因为它需要基于实际的工具使用、嘈杂的证据检索以及多约束推理。然而,大多数基准测试仅提供干净的地理空间数据和工具,缺少人们日常规划中使用的开放式社交信号。我们推出LifePlanner,这一基准测试用大规模本地社交媒体帖子丰富地图数据,并通过MCP工具集提供访问权限。LifePlanner提供涵盖四个任务类别和三个难度级别的评估套件。实验表明,前沿LLM在简单检索任务中表现良好,但在复杂规划任务中性能急剧下降,通过率降至40.2%。结果显示,失败主要源于从这个大型多模态数据库中获取的证据不完整、工具使用不精确以及约束整合能力弱,而非模型规模或推理长度,这表明未来的进展需要有效的基于实际的规划,而非单纯的规模扩大。
英文摘要
Geo-spatial planning, like trip design, is a realistic testbed for LLM agents because it requires grounded tool use, noisy evidence retrieval, and multi-constraint reasoning. Most benchmarks, however, only provide clean geospatial data and tools, missing the open-ended social signals that people use in daily planning. We introduce LifePlanner, a benchmark that enriches map data with large-scale local social media posts and provides access through an MCP toolset. LifePlanner provides an evaluation suite spanning four task categories and three difficulty levels. Experiments show frontier LLMs perform well on simple retrieval but degrade sharply on complex planning, with the Pass Rate dropping to 40.2%. Results show that failures mainly stem from incomplete evidence acquisition from such a large multimodal database, imprecise tool use, and weak constraint integration rather than model size or reasoning length, suggesting that future progress requires effective grounded planning instead of scaling alone.
Comments25 pages, 7 figures, 11 tables