arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度正则化JEPA世界模型从真实户外机器人数据中学习更多可转移表示

Depth-Regularized JEPA World Models Learn More Transferable Representations from Real Outdoor Robot Data

Usman M. Khan

arXiv 2607.16314首次发表:更新:

发表机构

Aigen(爱igen)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对从真实户外机器人数据学习世界模型的挑战,引入深度作为训练几何先验,结合SIGReg等方法,训练18M参数模型,实验表明该方法在降低视觉里程计误差、增加意外分数分离及提高多步展开保真度等方面有显著提升,增强了模型泛化能力。

AI 中文摘要

世界模型,尤其是基于JEPA架构的模型,已被证明能学习各种环境的鲁棒动力学。然而,从视觉复杂的现实世界数据学习仍是挑战,特别是在不可预测的户外环境。我们引入深度作为训练中的几何先验,直接从机器人视频数据学习更鲁棒的潜在动力学并处理视觉复杂性。这将深度监督与各向同性诱导潜在正则化器(SIGReg)结合,最大化与任务无关的潜在多样性并约束其组织方式,联合目标针对与场景几何一致的最高熵表示。为满足更高复杂性且不增加推理时间,还添加仅训练的过参数化。在真实农业机器人视频上训练一个18M参数模型,用冻结表示视觉里程计探针、基于预测器的意外检测和多步潜在展开保真度进行评估。与基线LeWM相比,我们的方法将视觉里程计探针误差降低33%,大幅增加域内和域外TartanGround基准上的意外分数分离,并改善域转移下的多步展开保真度,增益随展开视界增加。值得注意的是,在与3D几何无直接关联的物理理解(如光照和阴影)的意外分数分离上也有改进。这些结果表明,轻量级训练时几何先验使紧凑的JEPA世界模型在具有强大底层表示的真实户外数据上更有用且更可转移,而不增加推理开销。我们的工作表明,深度作为基于物理的先验可增强世界模型在各种任务上的泛化能力。

英文摘要

World models, especially based on JEPA architectures, have been shown to learn robust dynamics of various environments. However, learning from visually complex real-world data remains a challenge, especially in unpredictable outdoor environments. We introduce depth as a geometric prior during training in learning more robust latent dynamics directly from robot video data and handling visual complexity. This combines depth supervision with an isotropy-inducing latent regularizer (SIGReg), maximizing task-agnostic latent diversity while constraining how that diversity is organized, with the combined objective targeting the highest-entropy representation consistent with scene geometry. To satisfy this greater complexity without increasing inference time, we also add training-only overparameterization. Training an 18M-parameter model on video from a real agricultural robot, we evaluate with frozen-representation visual odometry probes, predictor-based surprise detection, and multi-step latent rollout fidelity. Compared to the baseline LeWM, our method lowers visual odometry probe error by 33%, substantially increases surprise-score separation both in-domain and on the out-of-domain TartanGround benchmark, and improves multi-step rollout fidelity under domain shift, with gains that grow with rollout horizon. Notably, we also see improvements in surprise-score separation on physics understanding that is not directly tied to 3D geometry, such as lighting and shadows. These results show that a lightweight training-time geometric prior makes a compact JEPA world model more useful and more transferable on real outdoor data with strong underlying representations, without adding inference overhead. Our work suggests that depth as a physically grounded prior can enhance world model generalization on a variety of tasks.

Comments13 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑