arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GRIP:高斯渲染作为图像到点云配准的跨模态桥梁

GRIP: Gaussian Rendering as a Cross-Modal Bridge for Image-to-Point Cloud Registration

Karim Slimani, Catherine Achard, Eric Marchand, Brahim Tamadazte

arXiv 2609.25966首次发表:更新:

发表机构

Inria; CNRS; INSERM; IRISA(法国国家信息与自动化研究所; 法国国家科学研究中心; 法国国家健康与医学研究院; 法国雷恩信息与系统研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GRIP通过高斯特征溅射将3D点特征渲染到图像网格,融合视觉与几何信息,实现图像到点云的精确配准,在RGB-D Scenes V2和7 Scenes上达到先进水平。

AI 中文摘要

本文介绍了GRIP,一个用于像素到点匹配和2D到3D配准的位姿条件细化框架。给定一个初始的粗略位姿估计,GRIP通过高斯特征溅射将学习到的3D点特征柔和地渲染到图像网格上,解决了基于网格的图像描述符与无序点云描述符之间的结构不匹配问题。渲染得到的点派生特征图随后通过像素对齐的Transformer与图像特征融合,使视觉语义和几何线索在共享的2D表示中交互。细化后的特征被解码并传播到更细的分辨率,用于密集对应估计和最终的位姿细化。在RGB-D Scenes V2和7 Scenes上的实验展示了最先进的内点比率和有竞争力的配准召回率,在更严格的评估阈值下表现更强。

英文摘要

This paper introduces GRIP, a pose-conditioned refinement framework for pixel-to-point matching and 2D to 3D registration. Given an initial coarse pose estimate, GRIP addresses the structural mismatch between grid based image descriptors and unordered point cloud descriptors by softly rendering learned 3D point features onto the image grid through Gaussian feature splatting. The rendered point derived feature map is then fused with image features by a pixel aligned transformer, enabling visual semantic and geometric cues to interact in a shared 2D representation. The refined features are decoded and propagated to finer resolutions for dense correspondence estimation and final pose refinement. Experiments on RGB D Scenes V2 and 7 Scenes demonstrate state of the art inlier ratio and competitive registration recall, with stronger performance under stricter evaluation thresholds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑