SVI2LoD3:基于智能体的、利用大语言与视觉模型从众包街景图像中重建语义三维城市模型的LoD3立面开口方法
SVI2LoD3: Agent-Driven Reconstruction of LoD3 Facade Openings in Semantic 3D City Models from Volunteered Street View Imagery using Large Language and Visual Models
浏览论文内容
中文总结 AI 辅助
SVI2LoD3提出一种由智能体驱动的零样本分割重建管道,结合新型FFD指标,可从街景图像生成符合标准的LoD3三维城市模型,减少标注量且评估更贴合实际需求
中文摘要 AI 辅助
本文提出了一种端到端、由智能体驱动的三维城市模型中立面开口的LoD3重建管道,可生成符合CityGML标准的直接可用输出。与依赖监督语义分割、因此需要大量人工标注训练数据的现有方法不同,所提方法采用零样本分割策略,大幅减少了标注工作量,同时在eTRIMS数据集基准测试中仍取得了优异性能。另一关键贡献是强制构建正确的部分层级结构,从而生成符合CityGML标准的LoD3建筑模型。除重建管道本身外,本研究还提出了一种名为立面特征距离(Facade Feature Distance,FFD)的新型立面重建评估指标。与主要通过像素级重叠评估相似性的传统指标(如mIoU或FRDS)不同,FFD通过视觉Transformer导出的高级特征空间测量距离,可同时捕捉语义正确性与建筑布局,为立面重建质量提供更合适的评估方式。所提管道与评估策略共同为语义丰富的三维城市模型的自动化生成与分析提供了实用且可扩展的贡献,开发的代码已发布于此:该URL
英文摘要
This paper presents an end-to-end, agent-driven pipeline for the LoD3 reconstruction of facade openings in 3D city models, producing directly usable CityGML-conform outputs. In contrast to existing approaches that rely on supervised semantic segmentation and therefore require large amounts of manually annotated training data, the proposed method employs a zero-shot segmentation strategy. This substantially reduces the annotation effort while still achieving strong performance in our benchmark on the eTRIMS dataset. A further key contribution is the enforcement of correct partonomic hierarchies, thereby producing CityGML-conform LoD3 building models. Beyond the reconstruction pipeline itself, this work also introduces a novel evaluation metric for facade reconstruction, termed Facade Feature Distance (FFD). Unlike conventional metrics such as mIoU or FRDS, which assess similarity primarily through pixel-wise overlap, FFD measures distance in a high-level feature space derived from a vision transformer. In doing so, it captures both semantic correctness and architectural layout, providing a more suitable assessment of facade reconstruction quality. The proposed pipeline and evaluation strategy together offer a practical and scalable contribution toward the automated generation and analysis of semantically enriched 3D city models. The developed code is published at: https://github.com/hcu-cml/citydb-SVI2LoD3-ai.
发表机构
- HafenCity University(汉堡港口城市大学)
机构由 AI 辅助整理,请以论文原文为准。