arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PARC-Loc:基于部分分配与关系一致性的文本到点云定位

PARC-Loc: Text-to-Point-Cloud Localization with Partial Assignment and Relational Consistency

Shengkai Ma, Zhenyu Hou, Weihua Cao

arXiv 2610.09761首次发表:更新:

AI 中文总结

PARC-Loc提出基于部分分配与关系一致性的由粗到精框架,解决文本到点云定位中布局不一致别名和边界证据不完整问题,在KITTI360Pose上Top-1召回率提升34%。

AI 中文摘要

文本到点云定位旨在根据周围物体的描述,在城市规模的3D地图中估计位置。现有的由粗到精方法利用聚合学习到的兼容性检索子地图,然后在选定的子地图内进行定位。然而,重复或相似的城市物体可能会夸大查询与多个子地图之间的嵌入相似性,即使子地图内的实例布局与查询描述不符。同时,与查询相关的实例往往跨越子地图边界,导致检索到的子地图上下文证据不完整。我们将这些失败模式分别称为布局不一致的别名和边界证据不完整。为解决这些问题,我们提出了PARC-Loc,一种基于部分分配与关系一致性(PARC)的由粗到精定位框架。PARC联合建模提示对象兼容性和成对空间关系,允许不匹配的元素,同时偏好与查询布局一致的分配。在粗阶段,其候选级评估补充了神经相似性,用于布局一致的子地图选择。在精阶段,上下文通过相邻子地图中与查询相关的实例进行扩展,而PARC产生对象级匹配权重,指导跨模态注意力。在KITTI360Pose和CityLoc上的大量实验表明,PARC-Loc优于传统的由粗到精基线。在KITTI360Pose上,我们的方法将5米处的Top-1定位召回率从0.50提高到0.67,相对于最强基线实现了34%的相对提升。

英文摘要

Text-to-point-cloud localization estimates a position in a city-scale 3D map from descriptions of surrounding objects. Existing coarse-to-fine methods retrieve submaps using aggregate learned compatibility and then localize within a selected submap. However, repetitive or similar urban objects can inflate the embedding similarity between the query and multiple submaps, even when the instance layout within a submap violates the query description. Meanwhile, query-relevant instances often span submap boundaries, leaving the retrieved submap with incomplete contextual evidence. We term these failure modes layout-inconsistent aliasing and boundary evidence incompleteness, respectively. To address them, we propose PARC-Loc, a coarse-to-fine localization framework built on Partial Assignment with Relational Consistency (PARC). PARC jointly models hint-object compatibility and pairwise spatial relations, allowing unmatched elements while favoring assignments consistent with the queried layout. At the coarse stage, its candidate-level assessment complements neural similarity for layout-consistent submap selection. At the fine stage, the context is expanded with query-relevant instances from adjacent submaps, while PARC yields object-level matching weights that guide cross-modal attention. Extensive experiments on KITTI360Pose and CityLoc show that PARC-Loc outperforms conventional coarse-to-fine baselines. On KITTI360Pose, our method improves Top-1 localization recall at 5 m from 0.50 to 0.67, achieving a 34% relative gain over the strongest baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑