arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20819cs.CV

4D基础模型能记住吗?

Can 4D Foundation Models Remember?

Guangzhao He, Hadar Averbuch-Elor, Wei-Chiu Ma

首次发表
浏览论文内容

中文总结 AI 辅助

针对4D基础模型视觉记忆能力不足的问题,提出PersistBench基准,利用360°视频作为真实标注,从物体恒存性、运动连续性和外观保持性三方面评估,揭示当前模型仅能维持短期一致性,为未来发展提供指导。

中文摘要 AI 辅助

感知和记忆视觉世界是导航和与环境交互的基础。当前的4D基础模型,如相机可控视频模型或4D重建模型,能够感知和重建动态环境,但它们对所感知内容的记忆能力如何仍是一个悬而未决的问题。现有基准主要依赖像素级指标,并且一旦物体离开视野就缺乏其真实标注,因此无法以物体为中心的方式对照参考来评估视觉记忆。为填补这一空白,我们引入了PersistBench,一个数据集和指标套件,利用360°视频作为全知真实标注,并提出三个评估方面:物体恒存性、运动连续性和外观保持性。对各种模型在多样类别上的评估显示,当前模型只能维持短期一致性,一旦物体离开视野,这种一致性就会显著下降。我们的发现凸显了当前模型能力与稳健视觉记忆之间的差距(“看见不等于记住”),为未来4D基础模型的发展提供了指导。数据集和代码可在项目页面上获取:此https URL。

英文摘要

Perceiving and remembering the visual world is fundamental to navigating and interacting with our environment. Current 4D foundation models, such as camera-controllable video models or 4D reconstruction models, can perceive and reconstruct dynamic environments, but how well they remember what they have perceived remains an open question. Existing benchmarks largely rely on pixel-level metrics and lack ground truth for objects once they leave the field of view, making them unable to evaluate visual memory in an object-centric manner against references. To fill this gap, we introduce PersistBench, a dataset and metric suite that leverages 360° videos as omniscient ground truth and proposes three evaluation aspects: object permanence, motion continuity, and appearance preservation. Evaluating various models across diverse categories reveals that current models can only maintain short-term consistency that degrades significantly once objects leave the field of view. Our findings highlight the gap between current model capabilities and robust visual memory ("seeing is not remembering"), providing guidance for future development of 4D foundation models. Dataset and code are available on the project page: https://guangzhaohe.com/persistbench.

发表机构

  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑