基于事件的多模态自我运动估计中学习表征的几何结构研究
On the Geometry of Learned Representations in Event-Based Multi-Modal Egomotion Estimation
- Politecnico di Milano(米兰理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究基于事件的多模态自我运动估计中多模态网络的几何结构,通过跨模态注意力架构融合多模态信号并训练,分析潜在空间几何和注意力动态,揭示相关特性,架起分析估计理论与数据驱动融合的桥梁。
AI中文摘要:
经典的基于事件的自我运动估计方法,包括ELOPE挑战赛中表现最佳的团队所采用的方法,依赖于几何优化框架,如对比度最大化、单应性估计或密集光流与解析运动反演相结合。本文研究了用于自我运动估计的多模态网络中出现的几何结构。通过跨模态注意力架构融合事件张量、惯性测量和距离信号,并在批处理设置中进行训练。分析了潜在空间几何和注意力动态,结果表明嵌入位于与运动变量对齐的低维流形上,注意力权重随角度激励和视觉可靠性而适应,融合表征恢复了经典的可观测性线索。这些结果架起了分析估计理论与现代数据驱动融合之间的桥梁。
英文摘要:
Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely on geometric optimization frameworks such as contrast maximization, homography estimation, or dense optical flow combined with analytic motion inversion. This work investigates the geometric structure that emerges inside a multi-modal network for egomotion estimation. Event tensors, inertial measurements, and range signals are fused through a cross-modal attention architecture and trained in a batch setting. We analyze the latent space geometry and attention dynamics, showing that (i) embeddings lie on low-dimensional manifolds aligned with motion variables, (ii) attention weights adapt with angular excitation and visual reliability, and (iii) the fused representation recovers classical observability cues. These results bridge analytical estimation theory and modern data-driven fusion.