arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从不确定性到确定性:无射线匹配的由粗到细视觉楼层平面图定位

From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching

Shiyong Meng, Bolei Chen, Ping Zhong, Yang Wan, Rongzhi Wang, Jiazhi Xia, Jianxin Wang

arXiv 2607.26817首次发表:更新:

AI 中文总结

该研究针对视觉楼层平面图定位的多模态姿态分布问题,提出无射线匹配的由粗到细框架,通过图像条件姿态扩散模型与局部细化器实现定位,在S3D和ZInD基准上达到最优精度与鲁棒性。

AI 中文摘要

视觉楼层平面图定位(FLoc)已成为室内定位的有前景解决方案,它通过将自我中心图像与极简结构地图匹配来实现定位。然而,由于跨模态信息不对称和重复的室内布局,视觉FLoc面临多模态姿态分布的根本挑战,即视觉上相同的观测结果会映射到不同且空间分离的位置。现有的基于射线匹配的方法通过显式预测稀疏几何或语义射线来应对这一问题,但这些方法固有地存在信息损失,且在推理过程中需要资源密集型的预处理和详尽的匹配。在本文中,我们绕过中间的射线匹配范式,提出了一种从不确定性到确定性的由粗到细视觉FLoc框架。在粗阶段,我们设计了一个图像条件姿态扩散模型来参数化连续的多模态姿态分布,有效地将随机初始化的姿态粒子引导到不同的候选模式。在细化阶段,我们提出了一个局部细化器,它从以候选为中心的楼层平面图裁剪中预测有界的亚米级姿态残差,其中结构歧义被大大消除。我们的方法在全局多假设跟踪和局部亚米级细化之间实现了有效平衡,无需任何离线地图预处理或测试时查找表。在S3D(完整)和ZInD基准上的综合结果表明,我们的方法达到了最先进的精度和鲁棒性。

英文摘要

Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist structural maps. However, due to cross-modal information asymmetry and repetitive indoor layouts, visual FLoc is fundamentally challenged by multimodal pose distributions, where visually identical observations map to distinct, spatially separated locations. Existing ray-matching-based methods tackle this by explicitly predicting sparse geometric or semantic rays, which inherently incur information loss and demand resource-intensive preprocessing alongside exhaustive matching during inference. In this paper, we bypass the intermediate ray-matching paradigm and propose a coarse-to-fine visual FLoc framework that progresses from uncertainty to determinism. In the coarse stage, we design an image-conditioned pose diffusion model to parameterize the continuous multimodal pose distribution, effectively routing stochastically initialized pose particles toward distinct candidate modes. In the refinement stage, we propose a localized refiner that predicts bounded sub-meter pose residuals from candidate-centered floorplan crops, where structural ambiguities are largely eliminated. Our method effectively balances global multi-hypothesis tracking and local sub-meter refinement without requiring any offline map preprocessing or test-time lookup tables. Comprehensive results on the S3D (full) and ZInD benchmarks demonstrate that our approach achieves state-of-the-art accuracy and robustness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑