发表机构
The Chinese University of Hong Kong, Shenzhen; Carnegie Mellon University(香港中文大学(深圳); 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RIDE利用重定位中的PnP-RANSAC内点对应关系生成稀疏度量深度,结合预训练视频深度模型先验,通过全局局部校正和时间记忆,实现无需微调的稠密深度估计,提升精度与时间一致性。
AI 中文摘要
渲染-匹配-PnP重定位通过建立查询图像像素与3D地图点之间的对应关系来恢复相机位姿,但其支持稠密深度估计的潜力常被忽视。为利用这一几何信息,我们提出RIDE,该方法从机器人的RGB流中估计稠密度量深度。给定度量缩放的3D高斯泼溅(3DGS)模型,RIDE将源自PnP-RANSAC内点对应关系的稀疏度量深度观测与预训练视频深度模型的几何先验相结合。为处理不均匀且间歇性的观测,它集成全局与局部深度校正及时间记忆,在度量尺度初始化后支持通过短观测间隙的深度估计。RIDE在公开RGB-D视频上训练,并在机器人序列上无需微调进行评估。实验表明,与仅尺度标定相比,深度精度和时间一致性均有提升,展示了定位几何如何同时支持位姿恢复和稠密机器人感知。
英文摘要
Render--match--PnP relocalization establishes correspondences between query image pixels and 3D map points for camera pose recovery, but their potential to support dense depth estimation is often overlooked. To exploit this geometric information, we present RIDE, which estimates dense metric depth from a robot's RGB stream. Given a metrically scaled 3D Gaussian Splatting (3DGS) model, RIDE combines sparse metric depth observations derived from PnP-RANSAC inlier correspondences with the geometric prior of a pretrained video-depth model. To handle uneven and intermittent observations, it integrates global and local depth correction with temporal memory, supporting depth estimation through short observation gaps after metric scale initialization. Trained on public RGB-D videos, RIDE is evaluated on robot sequences without fine tuning. Experiments show improved depth accuracy and temporal consistency over scale-only calibration, demonstrating how localization geometry can support both pose recovery and dense robot perception.