arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CityPlanner:一个用于可执行城市规划的沙盒智能体

CityPlanner: A Sandbox Agent for Executable Urban Planning

Wentao Zhang, Jingyuan Wang, Zetong Zhou, Yifan Yang, Wenrui Wang

arXiv 2609.09578首次发表:更新:

发表机构

School of Computer Science and Engineering, Beihang University; MIIT Key Laboratory of Data and Decision Intelligence, Beihang University(北京航空航天大学计算机科学与工程学院; 北京航空航天大学工业和信息化部数据与决策智能重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对城市规划中大规模行动空间下的决策难题,提出CityPlanner沙盒智能体框架,通过UrbanSandbox环境与原子任务强化学习分解规划流程,在真实基准上超越多种基线,验证了其有效性与可执行性。

AI 中文摘要

城市规划是一个现实世界的空间优化问题,需要在成本和服务质量等实际目标下,从庞大的候选空间中选取可行行动。现有的优化和强化学习方法对于固定公式化问题有效,但往往依赖于特定任务的表示和约束处理。我们提出了CityPlanner,一个用于可执行城市规划的沙盒智能体框架。CityPlanner引入了UrbanSandbox,一个统一的基于文件的环境,智能体在其中检查任务文件、生成计划、运行评估器,并根据可执行的反馈修订决策。为了使学习易于处理,我们进一步提出了原子任务强化学习,它将长的沙盒轨迹分解为用于初始构建的BuildPlan和用于基于反馈改进的ImprovePlan。在真实世界基准上的实验表明,CityPlanner始终优于启发式、特定任务强化学习和通用大语言模型智能体基线。消融实验验证了UrbanSandbox、原子任务强化学习和迭代部署的贡献。我们在该https URL发布了代码和数据集。

英文摘要

Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service quality. Existing optimization and reinforcement learning methods are effective for fixed formulations, but often depend on task-specific representations and constraint handling. We propose \emph{CityPlanner}, a sandbox-agent framework for executable urban planning. CityPlanner introduces \emph{UrbanSandbox}, a unified file-based environment where agents inspect task files, generate plans, run evaluators, and revise decisions based on executable feedback. To make learning tractable, we further propose atomic-task reinforcement learning, which decomposes long sandbox trajectories into \emph{BuildPlan} for initial construction and \emph{ImprovePlan} for feedback-based refinement. Experiments on a real-world benchmark show that CityPlanner consistently outperforms heuristic, task-specific RL, and general LLM-agent baselines. Ablations verify the contributions of UrbanSandbox, atomic-task RL, and iterative deployment. We release the code and dataset at https://anonymous.4open.science/r/co-agent-C1C8

CommentsEMNLP Under Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑