arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习应信任的对应关系:激光雷达地图中的置信度加权事件相机定位

Learning Which Correspondences to Trust: Confidence-Weighted Event-Camera Localization in LiDAR Maps

Panagiotis Kiousis, Kuangyi Chen, Jun Zhang, Friedrich Fraundorfer

arXiv 2610.11967首次发表:更新:

发表机构

Univ. of Patras; Institute of Visual Computing, Graz University of Technology(帕特雷大学; 格拉茨工业大学视觉计算研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对激光雷达地图中的事件相机定位问题,提出CELL方法,通过可微分概率PnP端到端学习对应关系置信度,结合部分补全深度表示,在M3ED和DSEC数据集上显著降低了中位平移和旋转误差。

AI 中文摘要

将事件相机在预先构建的激光雷达地图中定位,可转化为渲染深度图与事件图像之间的稠密光流估计,随后在生成的3D-2D对应关系上使用透视-n-点(PnP)求解器。现有流程在姿态估计期间依赖几何一致性,但未显式建模单个对应关系的可靠性或姿态信息量,即对应关系对相机姿态的约束强度。我们发现,通过逐对应关系误差约束置信度学习的自然方式存在深度相关偏差:小像素误差主要出现在大深度处,且不会带来高姿态信息量。相反,在我们的方法(CELL)中,我们通过可微分概率PnP的对数配分项,利用姿态端到端学习逐对应关系的置信度,该对数配分项鼓励能产生更受约束的姿态分布的权重配置。学习到的置信度有三种用途:(i)在解耦训练方案中对光流监督进行重加权,该方案使姿态梯度不进入光流/边缘骨干网络;(ii)在测试时驱动概率对应关系选择;(iii)与网络的边缘概率一起对最终的边缘匹配优化进行加权。我们还设计了部分补全深度表示,该表示能添加信号且不会在大间隙间产生幻觉。在M3ED和DSEC数据集上,我们的完整系统在大多数评估序列上优于LEAR基线:它将中位平移误差最多降低26.9%,中位旋转误差最多降低15.8%。

英文摘要

Localizing an event camera against a pre-built LiDAR map can be cast as dense optical-flow estimation between a rendered depth view and an event image, followed by a Perspective-n-Point (PnP) solver over the induced 3D-2D correspondences. Existing pipelines rely on geometric consensus during pose estimation, but do not explicitly model the reliability or pose informativeness, i.e., how strongly a correspondence constrains the camera pose, of individual correspondences. We show that the natural way to learn it -- using the per-correspondence error to constrain the learning of confidence -- suffers from a depth-dependent bias: small pixel errors reside predominantly at large depths and do not lead to high pose informativeness. Instead, in our method (CELL), we learn a per-correspondence confidence end-to-end through the pose, using a differentiable probabilistic PnP whose log-partition term encourages weight configurations that yield a better-constrained pose distribution. The learned confidence is used in three ways: (i) it reweights the flow supervision in a decoupled training scheme that keeps pose gradients out of the flow/edge backbone; (ii) it drives a probabilistic correspondence selection at test time; and (iii) together with the network's edge-probability it weights a final edge-matching refinement. We further design a partial-completion depth representation that adds signal without hallucinating across large gaps. On M3ED and DSEC our full system improves over the LEAR baseline on the majority of the evaluated sequences: it reduces the median translation error by up to 26.9% and the median rotation error by up to 15.8%.

Comments8 pages, 8 figures/tables. Submitted to IEEE ICRA 2027 (under review). Code: https://github.com/panagiotisq/CELL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑