arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05531cs.CVcs.RO

具身感知的深度先验学习

Deep Prior Learning for Embodied Perception

Yimou Wu, Jiaxin Guo, Yun-hui Liu, Zheng Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对具身感知中先验噪声与稀释问题,提出VPGGT框架,通过位姿损坏训练、先验残差连接和度量全局注意力,在四数据集上提升位姿精度。

中文摘要 AI 辅助

具身系统需要利用图像之外的可用观测的几何感知。最近的馈送式3D模型融入了几何先验,包括相机位姿、内参和深度。然而,处理噪声位姿、保持准确先验以及恢复物理尺度,需要的不仅仅是简单接受这些输入。我们引入了《视觉-先验几何接地变换器》(VPGGT),一个基于VGGT的框架,扩展了OmniVGGT以实现先验感知的具身感知。我们根据真实轨迹制定了传感器驱动的位姿损坏用于训练,并引入了一个无参数的《先验残差连接》(PRC)来缓解《先验稀释》,即预测比所提供的位姿先验更不准确的情况。我们的噪声公式针对相机位姿;提供的内参和深度不额外添加损坏。我们进一步引入了《度量全局注意力》,它根据可用的位姿和深度尺度调节一个全局尺度标记,并为几何输出预测一个共享的度量缩放因子。在四个数据集上的实验表明,当所有视图都提供相机先验时,在精确和损坏的位姿下,PRC相比匹配的训练基线提高了平移方向准确性和联合位姿AUC。这些结果支持在细化过程中显式访问先验作为特征级条件化的有用补充。

英文摘要

Embodied systems need geometric perception that exploits available observations beyond images alone. Recent feed-forward 3D models incorporate geometric priors, including camera poses, intrinsics, and depth. However, handling noisy poses, preserving accurate priors, and recovering physical scale require more than simply accepting these inputs. We introduce \emph{Vision-Prior Geometry Grounded Transformer} (VPGGT), a VGGT-based framework that extends OmniVGGT for prior-aware embodied perception. We formulate sensor-motivated pose corruptions from ground-truth trajectories for training and introduce a parameter-free \emph{prior residual connection} (PRC) to mitigate \emph{prior dilution}, where predictions are less accurate than their supplied pose priors. Our noise formulation targets camera poses; supplied intrinsics and depth receive no additional corruption. We further introduce \emph{Metric Global Attention}, which conditions a global scale token on available pose and depth scales and predicts a shared metric scaling factor for the geometric outputs. Experiments across four datasets show that \emph{PRC} improves translation-direction accuracy and joint pose AUC over a matched training baseline when camera priors are provided for all views, under both exact and corrupted poses. These results support explicit prior access during refinement as a useful addition to feature-level conditioning.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑