发表机构
The Remote Sensing Technology Institute; German Aerospace Center (DLR)(遥感技术研究所; 德国航空航天中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RoofDiT框架,以扩散Transformer和边预测模块为核心,结合几何感知注意力与多条件信号,在屋顶图合成与重建任务中取得优于基线的性能。
AI 中文摘要
我们提出了RoofDiT,这是一个用于二维屋顶图合成与重建的生成框架。屋顶被紧凑地描述为节点和结构边的平面图,但现有方法通常依赖固定的几何规则或直接重建目标。RoofDiT则直接将屋顶结构建模为顶点-边图,并学习其几何与连通性的条件生成先验。我们的框架采用两阶段设计:扩散Transformer生成屋顶顶点,边预测模块推断对应的图拓扑。为提升几何保真度,RoofDiT结合了感知相对几何的注意力机制与 footprint( footprint指建筑 footprint,即占地面积)和航拍图像条件,同时使用对齐正则化器以鼓励常见的水平、垂直和对角屋顶模式。通过改变条件信号,该模型可支持无条件生成、 footprint条件合成以及图像引导重建。实验表明,与扩散基线相比,图生成质量有所提升;在 footprint条件设置下,与直骨架先验相比表现更优;在图像引导重建中,达到了对比方法中最高的边F1值。
英文摘要
We present RoofDiT, a generative framework for 2D roof graph synthesis and reconstruction. Roofs are compactly described as planar graphs of junctions and structural edges, but existing methods often rely on fixed geometric rules or direct reconstruction objectives. RoofDiT instead models roof structures directly as vertex-edge graphs and learns a conditional generative prior over their geometry and connectivity. Our framework follows a two-stage design: a diffusion transformer generates roof vertices, and an edge prediction module infers the corresponding graph topology. To improve geometric fidelity, RoofDiT combines relative geometry-aware attention with footprint and aerial-image conditioning, while using an alignment regularizer to encourage common horizontal, vertical, and diagonal roof patterns. The same model supports unconditional generation, footprint-conditioned synthesis, and image-guided reconstruction by changing the conditioning signal. Experiments show improved graph generation quality over a diffusion baseline, favorable performance against a straight-skeleton prior in the footprint-conditioned setting, and the highest edge F1 among compared methods for image-guided reconstruction.
CommentsPre-review manuscript. Accepted at the ICPR 2026 Workshop on Pattern Recognition in Remote Sensing (PRRS)