arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30056cs.ROcs.CV

M3GD:用于相机-激光雷达新视角合成的多模态多视角几何扩散

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

  • New York University(纽约大学)
  • University of California, Berkeley(加州大学伯克利分校)
  • U.S. Army Combat Capabilities Development Command, Army Research Laboratory(美国陆军作战能力发展司令部陆军研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, Giuseppe Loianno

AI总结:

M3GD利用相机-激光雷达多模态几何扩散,通过共享空间结构将激光雷达特征注入图像生成器,在GrandTour数据集上提升新视角RGB和深度合成质量,并支持地面机器人部署。

AI中文摘要:

机器人新视角合成(NVS)必须同时恢复视觉外观和度量三维结构,然而大多数生成式NVS方法仅依赖图像,忽视了激光雷达——一种在机器人平台上常见的互补传感器。我们提出M3GD,一种用于生成式NVS的相机-激光雷达多模态表示,它组合了独立预训练的2D图像和3D点云基础模型,无需单独预训练跨模态转换器。我们证明,在相机投影后,冻结的激光雷达和图像特征展现出显著共享的空间结构,提供了一种自然的跨模态表示。M3GD通过这种结构以激光雷达为条件进行生成:它将显式几何统计与学习到的点云描述符组合成图像潜在网格上的视图对齐数据包,通过轻量级残差适配器注入多视图流匹配生成器,该生成器的潜在空间、解码器和训练目标保持不变。在GrandTour数据集上,M3GD在目标视图RGB和深度合成方面优于同一骨干网络的仅图像版本。消融实验表明,性能提升来自像素对齐的激光雷达内容,且目标视图激光雷达充当几何查询,将请求的视图与源观测联系起来。在地面机器人上的部署展示了实际的现实世界操作,通过欧拉积分步数控制可配置的质量-成本权衡。

英文摘要:

Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator. We show that, after camera projection, frozen LiDAR and image features exhibit substantial shared spatial structure, providing a natural cross-modal representation. M3GD conditions generation on LiDAR through this structure: it combines explicit geometry statistics with learned point-cloud descriptors into view-aligned packets on the image-latent grid, injected through a lightweight residual adapter into a multi-view flow-matching generator whose latent space, decoders, and training objective remain intact. On the GrandTour dataset, M3GD improves target-view RGB and depth synthesis over an image-only version of the same backbone. Ablations show that the gains come from pixel-aligned LiDAR content and that target-view LiDAR acts as a geometric query linking the requested view to source observations. Deployment on a ground robot demonstrates practical real-world operation, with a configurable quality--cost trade-off controlled by the number of Euler integration steps.

↑