DUDA-Bench:多模态数据驱动城市诊断的LLM智能体基准
DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出DUDA-Bench基准,将多模态城市诊断形式化为多阶段智能体工作流,评估发现端到端诊断与孤立分析能力差距显著,并揭示跨阶段协调的关键局限。
AI中文摘要:
城市诊断整合异构观测数据,以识别城市问题、定位受影响区域并调查促成因素,为循证城市规划和治理提供信息。然而,其依赖劳动密集且针对特定案例的专家工作流,限制了可扩展性和复用性,从而推动了对基于智能体执行的探索。为评估这一能力,我们引入了DUDA-Bench,一个分层且交互式的基准,将数据驱动的城市诊断形式化为多阶段智能体工作流。该基准包含86个原子任务和22个工作流任务,涵盖四个分析阶段,基于来自12个城市的多模态数据,覆盖五种城市问题类型。对七个骨干模型和五个智能体系统的评估显示,孤立分析能力与端到端诊断之间存在显著差距,且系统收益因骨干模型而异。轨迹分析表明,未解决的证据缺口会在各阶段间传播,而成功的恢复涉及利用反馈修正假设和行动。这些发现凸显了跨阶段协调分析能力的局限性,特别是自适应规划、证据整合和验证方面。更广泛地,DUDA-Bench提供了一个框架,将专家分析工作流转化为分层智能体任务和过程感知评估,支持对端到端分析能力的系统性评估。
英文摘要:
Urban diagnosis integrates heterogeneous observations to identify urban problems, localize affected areas, and investigate contributing factors, informing evidence-based urban planning and management. However, its reliance on labor-intensive, case-specific expert workflows limits scalability and reuse, motivating the exploration of agent-based execution. To evaluate this capability, we introduce DUDA-Bench, a hierarchical and interactive benchmark that formalizes data-driven urban diagnosis as a multi-stage agent workflow. It comprises 86 atomic and 22 workflow tasks spanning four analytical stages, grounded in multimodal data from 12 cities covering five urban problem types. Evaluations of seven backbone models and five agent systems reveal a substantial gap between isolated analytical competence and end-to-end diagnosis, with system benefits varying across backbones. Trajectory analysis shows that unresolved evidence gaps propagate across stages, while successful recovery involves revising assumptions and actions using feedback. These findings highlight limitations in coordinating analytical capabilities across stages, particularly adaptive planning, evidence integration, and verification. More broadly, DUDA-Bench provides a framework for translating expert analytical workflows into hierarchical agent tasks and process-aware evaluation, supporting systematic assessment of end-to-end analytical capabilities.