发表机构
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对城市干预评估昂贵的问题,提出多智能体系统VIDA-Geo,结合分割、扩散修复和指标评分生成干预方案,在8项指标上优于基线,支持规划师参与决策。
AI 中文摘要
城市环境由设计选择塑造,这些选择对健康、安全和生活质量具有长期影响,然而评估拟议的干预措施仍然成本高昂、耗时且往往不切实际。现有的地理空间视觉方法主要侧重于从航空和街景图像中监测城市指标,而非提出干预措施并估计其对这类指标的影响。超越识别范畴,我们引入了为给定航空或街景图像发现可改善目标指标的干预措施这一问题。我们认为,黑盒指标模型与生成式编辑模型相结合,可以充当测试干预假设的隐式数字孪生。我们提出了VIDA-Geo,一个多智能体系统,通过协调分割、基于扩散的图像修复和指标评分模型来探索这一干预空间,以产生既在感知上逼真又与真实世界政策相符的干预措施。我们在航空和街景图像上的8项指标上评估了我们的系统,测量了感知安全性和绿化等因素的变化。我们的方法在许多情况下优于现有基线,实现了高达2倍的感知质量和政策一致性得分。最后,我们的模型为用户提供多个候选干预措施,支持专家城市规划师参与的工作流程。
英文摘要
Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventions remains costly, time-consuming, and often impractical. Existing geospatial vision methods largely focus on monitoring urban indicators from aerial and street-view imagery, rather than proposing interventions and estimating their effects on such indicators. Moving beyond recognition, we introduce the problem of discovering interventions that improve target indicators for a given aerial or street-view image. We argue that a black-box indicator model, combined with a generative editing model, can serve as an implicit digital twin for testing intervention hypotheses. We present VIDA-Geo , a multi-agent system that explores this intervention space by coordinating segmentation, diffusion-based inpainting, and indicator scoring models to produce interventions that are both perceptually realistic and aligned with real-world policies. We evaluate our system on 8 indicators across aerial and street-view imagery, measuring changes in factors such as perceived safety and greenery. Our approach outperforms existing baselines in many cases, achieving up to 2X higher perceptual quality and policy alignment scores. Finally, our model provides users with multiple candidate interventions, supporting an expert city-planner-in-the-loop workflow.