arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03137cs.AI

当帧中很少物体移动时,保持JEPA世界模型的可规划性

Keeping JEPA World Models Plannable When Little of the Frame Moves

Florian Strohm, Patrick Wagner, Jannik Schwab, Marco Huber

首次发表
浏览论文内容

中文总结 AI 辅助

针对JEPA世界模型在帧中少物体移动时规划失败的问题,提出逆动力学辅助损失修复编码器动作敏感性,使语言目标规划成功率从0.003提升至0.35,并验证了动作敏感性探测作为规划必要条件。

中文摘要 AI 辅助

以语言而非目标帧来指定目标,是与潜在世界模型进行规划的自然接口,但对其进行测试需要场景中语言必须区分多个物体。我们构建了SLIM,一个包含多个小物体以及相同场景上配对的视觉和语言目标的推动基准。在SLIM上,解决PushT的LeWM世界模型在不到1%的试验中成功,而一个具有模拟器状态的脚本控制器则解决了每个层级。探测将失败定位在编码器上:其潜在表示几乎对动作不敏感,无法从中解码推杆或物体位置,并且展开不比将当前潜在表示向前复制更好。一个逆动力学辅助损失,应用于编码器潜在表示和通过共享头预测的潜在表示(在测试时丢弃),恢复了每个探测并将成功率从0.003提高到0.35(在困难推动层级上为0.16,其中目标无关策略得分为零),并在两倍训练视界上改进了PushT。控制实验将修复归因于进入编码器的梯度,响应扫描显示,当帧中有足够部分对动作响应时,原始模型能够规划。一个廉价的动作敏感性探测,无需环境访问即可计算,作为经验必要条件:所有低于其阈值的配置都未能规划。在修复后的潜在表示上,一个小型语言目标头无需重新训练世界模型即可从句子进行规划:在导航上达到0.84(视觉目标神谕为1.00),当命名区域与诱饵交换时跟随该区域,并对未见名词优雅退化。单个目标句子很少完成推动,但给定推动作为阶段句子的序列,该头将中等和困难推动层级的成功率从0.04提高到0.25,与目标帧神谕相当,即使阶段之间的切换仅从潜在表示读取也是如此。

英文摘要

Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in which language must discriminate between several objects. We build SLIM, a pushing benchmark with several small objects and paired visual and language goals on identical scenes. On SLIM a LeWM world model that solves PushT succeeds on under 1% of trials, although a scripted controller with simulator state solves every tier. Probes locate the failure in the encoder: its latent is nearly action-insensitive, neither pusher nor object positions can be decoded from it, and rollouts are no better than copying the current latent forward. One inverse-dynamics auxiliary loss, applied to encoder latents and to predicted latents through a shared head discarded at test time, restores every probe and raises success from 0.003 to 0.35 (0.16 on the hard pushing tier, where a goal-agnostic policy scores zero), and improves PushT at twice the trained horizon. Controls attribute the repair to the gradient into the encoder, and a response sweep shows that the vanilla model plans once enough of the frame responds to actions. A cheap action-sensitivity probe, computable without environment access, acts as an empirical necessary condition: all configurations below its threshold failed to plan. On the repaired latent, a small language-goal head plans from sentences without retraining the world model: it reaches 0.84 on navigation (visual-goal oracle 1.00), follows the named zone when it is swapped with a decoy, and degrades gracefully to unseen nouns. A single goal sentence rarely completes a push, but given the push as a sequence of stage sentences the head raises success on the medium and hard pushing tiers from 0.04 to 0.25, on par with the goal-frame oracle, also when the switch between stages is read from the latent alone.

发表机构

  • Fraunhofer IPA(弗劳恩霍夫制造工程与自动化研究所)
  • University of Stuttgart(斯图加特大学)

机构由 AI 辅助整理,请以论文原文为准。

↑