arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于多模态大语言模型(MLLM)的遥感城市布局提取

Remote-Sensing City Layout Extraction with MLLM

Zigan Zhou, Kai Li, Yupeng Deng

arXiv 2608.16484首次发表:更新:

发表机构

City University of Hong Kong; University of Chinese Academy of Sciences; Aerospace Information Research Institute(香港城市大学; 中国科学院大学; 空天信息创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Code-as-City方法,利用MLLM将遥感顶视图图像转化为含平面与3D输出的可编辑城市布局,在CityLayout-100数据集上取得41.1%、48.3%的交并比,验证了视觉观测转城市代码的可行性。

AI 中文摘要

遥感系统通常用检测框、语义掩码或矢量边界描述城市内容,这类输出能定位类别并支持图像平面评分,但本身无法构成保留对象标识、类型关系、拓扑结构和再生规则的可执行布局。Code-as-City将从单一顶视图图像提取城市布局的任务,转化为用多模态大语言模型(MLLM)实现的约束代码生成任务。图像模型首先生成对齐的五类语义布局先验,三次有序的MLLM传递会结合该图像与先验,恢复道路、土地覆盖区域及其关系、建筑物。确定性归一化将累积的记录转换为城市图和受限布局程序,执行该程序可生成可渲染的3D城市布局,以及在共享几何上的正交语义投影;该投影可与遥感掩码进行像素级比较,同时命名对象、关系和编辑操作仍可用于同步再生两种视图。在CityLayout-100的100个场景上评估,整个框架获得41.1%的平均交并比和48.3%的全局交并比,该结果提供了定量证据,表明视觉观测可转化为具有耦合平面和3D输出的可检查、可编辑城市代码。

英文摘要

Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector boundaries. Such outputs locate classes and support image-plane scoring, yet they do not by themselves constitute an executable layout that retains object identities, typed relations, topology, and regeneration rules. Code-as-City instead casts urban-layout extraction from a single top-down image as constrained code generation with a multimodal large language model (MLLM). An image model first produces an aligned five-class semantic layout prior. Three ordered MLLM passes use the image and this prior to recover roads, land-cover regions and relations, and buildings. Deterministic normalization converts the accumulated records into a city graph and a restricted layout program. Executing the program creates a renderable 3D city layout and an orthographic semantic projection over shared geometry. The projection admits pixel-level comparison with remote-sensing masks, while named objects, relations, and editing operations remain available for synchronized regeneration of both views. Evaluated on the 100 scenes of CityLayout-100, the complete framework obtains 41.1% mean intersection-over-union and 48.3% global intersection-over-union. This result provides quantitative evidence that visual observations can be translated into inspectable, editable city code with coupled planar and 3D outputs.

Comments4 pages, 2 figures, 4 tables. Accepted to IEEE APGARSS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑