arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潜在规划在点云场景中是否可行?面向几何观测的动作条件JEPA世界模型

Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals

Fabio F. Oberweger, Michael Schwingshackl, Markus Murschitz

arXiv 2608.29434首次发表:更新:

发表机构

AIT Austrian Institute of Technology; Assistive & Autonomous Systems(奥地利技术研究所; 辅助与自主系统)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究将JEPA模型扩展至点云场景,验证了三种JEPA设计可在点云规划中有效工作,其中动作敏感型模型表现最优,还实现了无需目标观测的3D目标接口构建。

AI 中文摘要

JEPA世界模型使潜在空间规划成为控制领域的实用路径,但这类模型几乎仅基于图像构建。潜在预测是否适用于几何观测尚不明确:点云具有稀疏、无序和自遮挡的特性,且场景中0.3%-15%的点会发生移动,潜在预测的慢特征最优性会与3D自监督的几何捷径效应叠加。我们将三种经典JEPA设计(冻结编码器、分布先验、动作敏感型)扩展至点云场景,并重新设计stable-worldmodel基准,使其仅观测方式与图像基准不同。三种模型均能实现规划且未出现崩溃:分布先验模型在所有基准上的表现与重新评估的图像模型统计等价,动作敏感型模型在几何移动最多的受控对比中取得最优结果。探测分析揭示了原因:物体位置几乎可完美线性解码,注意力集中在少数移动点上。规划可承受训练中从未出现的大量 dropout,但范围噪声会破坏最稀疏的场景。几何观测最终使指定的3D目标成为自然的目标接口:我们从目标和当前潜在中构建目标潜在,成功率未受影响,且无需目标观测。

英文摘要

Latent action world models let agents plan new behaviors at test time by predicting how actions change the environment, and joint-embedding predictive architectures (JEPAs) do so by forecasting future latent states rather than pixels. Yet nearly all such models see the world through a camera, even though robotic manipulation is fundamentally geometric: in robotics goals for manipulation are traditionally specified by target object poses, not by images of the object once placed. We ask whether latent planning survives a shift from appearance to geometry, on the observation side as well as on the goal specifications side. To answer this, we extend the stable-worldmodel evaluation platform with simulated LiDAR-style raycast point clouds as a new sensor modality, and adapt three JEPA designs to point clouds: a frozen-encoder model built on Utonia features, a distribution-prior model based on LeWM, and an action-sensitive model based on Delta-JEPA. We further introduce a goal-encoding mechanism that constructs the goal latent from the current latent and a 3D target pose, removing the need for goal images or goal point clouds. A comparative evaluation of the different anti-collapse mechanisms shows that point-cloud world models can match their image-based counterparts, demonstrating that the modality shift from appearance to geometry is achievable. All models are released as open weights with open-source training and inference code, to make world-model planning accessible for LiDAR-driven and pose-directed robotic tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑