arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DAP-Pose:用于鲁棒姿态估计的深度时间对齐和物理感知跨模态传感器融合

DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation

Jianhan Lin, Yuchu Qin, Jiateng Yuan, Wenbo Zhang, Shuai Gao

arXiv 2607.23755首次发表:更新:

发表机构

Aerospace Information Research Institute, Chinese Academy of Sciences; International Research Center of Big Data for Sustainable Development Goals; University of Chinese Academy of Sciences; The University of Adelaide(中国科学院空天信息创新研究院; 可持续发展大数据国际研究中心; 中国科学院大学; 阿德莱德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对复杂环境下多模态传感器的姿态估计问题,提出DAP-Pose模型,通过双级跨模态融合模块捕捉线索,深度时间对齐模块处理异步流,结合物理感知约束,在KITTI数据集上达最优性能,平均平移误差1.31%,旋转误差0.46°,严重错位下也能准确估计。

AI 中文摘要

在复杂环境中,使用多模态传感器进行鲁棒且准确的姿态估计对于自动驾驶车辆和移动机器人系统至关重要。本文提出了DAP-Pose,一个用于鲁棒多模态姿态估计的统一端到端模型。它引入了双级跨模态融合(BCF)模块,从视觉、惯性和GNSS测量中捕捉互补的语义和几何运动线索。设计了深度时间对齐(DTA)模块处理时间偏移,在潜在空间中显式对齐异步流。还通过流形几何和GNSS引导的绝对度量尺度纳入物理感知约束。在公共KITTI基准数据集上实验表明,DAP-Pose达到了最优性能,平均平移误差最低为1.31%,旋转误差为0.46°,且在严重时间错位下也能准确估计姿态并保持鲁棒性能。

英文摘要

Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ($t_{rel}$) of 1.31% and rotation error ($r_{rel}$) of 0.46$^{\circ}$. Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑