arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03018cs.AI

UrbanAgent:面向跨系统城市任务的工具增强型智能体

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

Jiayu Cao, Xingyuan Zeng, feiyu Li, Zhijing Huang, Xujie Yuan, Rongxiang Chen, Shimin Di, Libin Zheng, Jian Yin

首次发表
浏览论文内容

中文总结 AI 辅助

针对城市数字服务碎片化问题,提出工具增强型智能体UrbanAgent,结合大语言模型与多类工具,引入基准Urban-Eval评估,在多模型上任务成功率达71%,领先基线10个百分点。

中文摘要 AI 辅助

现代城市运营依赖的数字服务数量日益增多,但居民日常需求仍难以满足,这些服务碎片化且互操作性差,给用户带来沉重操作负担。现有数字平台、城市基础模型和智能助手仅能解决城市任务的孤立方面,难以将复杂自然语言请求可靠转化为可执行的跨系统工作流。我们提出UrbanAgent,一种面向跨系统城市任务的工具增强型智能体框架,它结合大语言模型的认知与推理能力,以及支持代码执行、API调用和模型上下文协议(Model Context Protocol)的工具集。通过自适应闭环,它在行动前明确缺失信息,基于实时观测确定工具使用,并使最终响应与观测证据及任务约束保持一致。为解决评估缺口,我们引入Urban-Eval,一种专为跨系统城市请求设计的基准,不同于以往仅评估通用工具使用或城市知识与推理的基准,Urban-Eval同时评估任务结果与执行质量,包括所需工具覆盖率、依赖有效性和证据可追溯性。实验结果表明,UrbanAgent任务成功率达71%,比最强基线高出10个百分点,该优势在GPT-5-mini、Gemini-2.5-flash、DeepSeek-V4-flash和Qwen3-235B-A22B模型上均成立。

英文摘要

Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragmented and have little interoperability, placing a heavy operational burden on users. Existing digital platforms, urban foundation models, and intelligent assistants each address only isolated aspects of an urban task. But they struggle to reliably convert complex natural-language requests into executable cross-system workflows. We propose Urban-Agent, a tool-augmented agent framework for cross-system urban tasks. It couples the cognitive and reasoning capabilities of a large language model with a tool-set supporting code execution, API calls, and Model Context Protocol. Through one adaptive closed loop, it clarifies missing information before acting, grounds tool use in live observations, and aligns the final response with observed evidence and task constraints. To address the evaluation gap, we introduce Urban-Eval, a benchmark specifically designed for cross-system urban request. Unlike prior benchmarks that assess either general tool use or urban knowledge and reasoning, Urban-Eval evaluates both task results and execution quality, including required tool coverage, dependency validity, and evidence traceability. Experimental results indicate that Urban-Agent reaches a 71% task success rate, 10 points above the strongest baseline. This lead holds across GPT-5-mini, Gemini-2.5-flash, DeepSeek-V4-flash, and Qwen3-235B-A22B.

↑