arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于RNN和注意力机制的端到端视觉里程计

End-to-End Visual Odometry with RNNs and Attention

Ruiyu Li, Yinjia Liu, Alexander Yu

arXiv 2609.26188首次发表:更新:

AI 中文总结

本文研究端到端深度学习视觉里程计,提出基于时间注意力的模型改进基线,并探索其在手持摄像头动态场景下的性能。

AI 中文摘要

视频里程计(VO)是通过分析视觉信息(如来自一个或多个摄像头的帧序列)来估计物体自身运动的过程。它一直是计算机视觉和机器人学中的热门研究课题,其应用包括移动机器人系统以及自动驾驶。在本项目中,我们研究了现有的端到端深度学习VO方法,并提出了一种新颖的基于时间注意力的模型来改进基线。此外,尽管现有绝大多数基于深度学习的VO方法都是在驾驶数据上训练的,但我们研究了基于深度学习的VO在更具动态性和复杂性的手持摄像头问题上的性能。

英文摘要

Video Odometry (VO) is the process of estimating the ego-motion of an object by analyzing visual information such as a sequence of frames from one or multiple cameras. It has been a popular research topic in computer vision and robotics, and its applications include mobile robotic systems as well as autonomous driving. In this project, we investigate existing end-to-end deep-learning approaches to VO, and propose a novel temporal attention-based model to improve upon the baseline. In addition, while the vast majority of existing deep-learning-based approaches to VO are trained on driving data, we investigate the performance of deep-learning-based VO to the more dynamic and complex problem of hand-held cameras.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑