多模态平面图编码:学习密集模态不变表示
Multimodal Floorplan Encoding: Learning Dense Modality-Invariant Representations
- University of Zaragoza(萨拉戈萨大学)
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出多模态平面图编码器(MMFE),通过共享密集潜在网格和InfoNCE损失对齐跨模态表示,结合几何一致性增强,在Structured3D上实现跨模态匹配、对齐和检索的显著提升。
AI中文摘要:
平面图以多种形式出现,从矢量CAD图纸到光栅渲染和传感器导出的密度图。这种异质性使得构建能够跨模态迁移并支持几何中心任务(如对齐和检索)的学习系统变得困难。我们引入了多模态平面图编码器(MMFE),它将多样的2D室内表示映射到一个共享的密集潜在网格中。MMFE结合了冻结的DINOv3骨干网络和一个可训练的密集预测Transformer(DPT)头,并通过逐细胞的信息噪声对比估计(InfoNCE)目标进行训练,该目标在跨模态中对齐空间对应的区域,同时使用所有其他细胞作为负样本。为了提高对几何失真的鲁棒性,我们引入了受控的相似变换,并通过特征网格扭曲来强制几何一致性。在Structured3D(一个保留的域外数据集)上,MMFE改善了跨模态密集匹配,实现了与RANSAC的鲁棒相似性对齐,并在与学习聚合相结合时产生了强大的检索性能。
英文摘要:
Floorplans arise in many forms, from vector CAD drawings to raster renderings and sensor-derived density maps. This heterogeneity makes it difficult to build learning systems that transfer across modalities and support geometry-centric tasks such as alignment and retrieval. We introduce the Multimodal Floorplan Encoder (MMFE), which maps diverse 2D indoor representations into a shared dense latent grid. MMFE combines a frozen DINOv3 backbone with a trainable Dense Prediction Transformer (DPT) head, and is trained with a per-cell Information Noise-Contrastive Estimation (InfoNCE) objective that aligns spatially corresponding regions across modalities while using all other cells as negatives. To improve robustness to geometric distortions, we incorporate controlled similarity transformations and enforce geometric consistency through feature-grid warping. On Structured3D, a held-out out-of-domain dataset, MMFE improves cross-modal dense matching, enables robust similarity alignment with RANSAC, and yields strong retrieval when paired with learned aggregation.