arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

追踪并非持久性:视频世界模型对隐藏物体保留了什么

Tracking Is Not Permanence: What Video World Models Keep of a Hidden Object

Peng Xie, Amr Alanwar

arXiv 2610.07355首次发表:更新:

发表机构

Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨视频世界模型对隐藏物体的追踪持久性,发现预测器快速丢失隐藏物体,而编码器保留信息,且可通过训练低成本植入持久性先验。

AI 中文摘要

视频世界模型能够追踪它们可见的物体;我们探究它们对无法看到的物体保留了什么。我们将一个物体从冻结的V-JEPA 2预测器中隐藏起来,并将其对隐藏区域的预测与编码器对仅在区域内不同的两个世界的表示进行比较。预测器的决策部分保留了一个静止物体,而完全未保留容器内携带的物体,并在0.3秒内丢失一个移动物体(在V-JEPA自身的管状掩码下为0.5秒;ViT-H在预训练90%掩码率下将其保持至1.1秒);在投影中,痕迹保留在中点以下,为复制最后视图的基线所保留量的14-60%。信息确实存在:编码器以1.00读取物体的存在,并保持封闭容器内容可解码3.5秒,而预测器的输出,用编码器自身的探针读取,在盒子关闭半秒后,在2%的场景中包含球。在渲染场景中,预测器一侧缺乏持久性,而训练将其作为先验廉价地安装:在合成容器上进行三千步仅预测器训练,将此信念从0.05提升至1.00,而两个匹配对照则不然。它们还将IntPhys-2019从84.2%提升至93.3%,但没有容器的课程也如此,并且基准所归功的训练习惯随其评分规则而变化。使用管状掩码的持续训练在操作和互联网风格视频上产生1.1-1.6秒的移动物体携带效应,因此该缺陷并非潜在预测所固有。VideoMAE几乎不保留任何内容,而Cosmos的下一令牌预测保留了一个静止的隐藏物体,但不保留移动容器内携带的物体。

英文摘要

Video world models track objects they can see; we ask what they keep of objects they cannot. We hide an object from a frozen V-JEPA 2 predictor and compare its prediction for the hidden region with the encoder's representation of two worlds that differ only inside that region. The predictor's decision keeps a stationary object in part and one carried inside a container not at all, and loses a moving one within 0.3 s (0.5 s under V-JEPA's own tube mask; ViT-H keeps it to 1.1 s at pretraining's 90% masking ratio); in projection a trace remains, below the midpoint, at 14-60% of what a baseline copying the last view retains. The information is there: the encoder reads the object's presence at 1.00 and keeps a closed container's contents decodable for 3.5 s, while the predictor's output, read with the encoder's own probe, contains the ball in 2% of scenes once the box has been closed for half a second. On rendered scenes, permanence is missing on the predictor's side, and training installs it cheaply as a prior: three thousand predictor-only steps on synthetic containers take this belief from 0.05 to 1.00 against two matched controls. They also raise IntPhys-2019 from 84.2% to 93.3%, but so does a curriculum without containers, and which training habit the benchmark credits changes with its scoring rule. Continued training with tube masks produces 1.1-1.6 s of moving-object carry-over on manipulation and internet-style video, so the deficit is not intrinsic to latent prediction. VideoMAE keeps almost nothing, and Cosmos's next-token prediction keeps a stationary hidden object but not one carried inside a moving container.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑