arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28414cs.CVcs.AIcs.RO

冻结的流会遗忘:诊断并恢复潜在流世界模型中的丢失运动

Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model

Xiwen Chen, Rigaudiere Z. Li, Zhiruo Zhou, Xiaojun Zhu, Houde Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对冻结潜在流世界模型中运动丢失问题,提出解码增强的展开训练(DART),仅重训练流并利用解码路径监督,恢复运动结构并提升预测质量,缩小与理想插值的差距。

中文摘要 AI 辅助

在冻结的自监督潜在空间中集成流的潜在世界模型训练稳定且廉价,但会静默丢失属性操作最依赖的运动。预训练流从不移动被操作对象;仅使用潜在损失重新训练它只会将静止转换为类似瞬移的运动。我们将失败追溯到训练信号,而非表示:锚点稀疏、仅潜在监督从未说明变化在时间范围中的位置。解码增强的展开训练(DART)在保持表示冻结的同时修复了这一点,仅使用解码路径监督重新训练流。DART在完整协议上优于其仅潜在父模型,恢复了运动的时间结构,并将预测运动重新耦合到场景;在更大规模下,它进一步提高了预测质量,缩小了与一个预言式插值参考之间近一半的剩余差距。最后,我们报告了一个关于评估的意外发现:像素误差单独会奖励冻结的预测。

英文摘要

Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow never moves the manipulated object; retraining it with latent-only losses only trades stillness for teleport-like motion. We trace the failure to the training signal, not the representation: anchor-sparse, latent-only supervision never says where along the horizon change belongs. Decode-augmented rollout training (DART) repairs this while keeping the representation frozen, retraining only the flow with decode-path supervision. DART outperforms its latent only parent on the full protocol, restores the temporal structure of motion, and re-couples predicted motion to the scene; at larger scale it further improves prediction quality, closing nearly half the remaining gap to an oracle-informed interpolation reference. Finally, we report an unexpected finding about evaluation: pixel error alone rewards frozen predictions.

发表机构

  • Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
  • Shanghai Jiaotong University(上海交通大学)
  • Wuhan University of Technology(武汉理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑