OPIS:视频世界模型中多对象记忆的输入锚定基准
OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models
- USTC(中国科学技术大学)
- CASIA(中国科学院自动化研究所)
- HKUST(香港科技大学)
- Infrec
- Tsinghua University(清华大学)
- UCAS(中国科学院大学)
- The Chinese University of Hong Kong(香港中文大学)
- Sun Yat-sen University(中山大学)
- University of Waterloo(滑铁卢大学)
- Jiangnan University(江南大学)
- NUS(新加坡国立大学)
- Fiveages
- SUTD(新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
OPIS是一个输入锚定的基准,用于评估视频世界模型的多对象记忆,通过对象中心评估器分层测量存在、身份和结构,发现保留特定对象实例比生成合理视觉元素更难。
AI中文摘要:
视频世界模型必须随时间保持世界的视觉状态,但现有的评估协议往往依赖于生成的轨迹、视频参考或选定的重访视角,这些可能会混淆对模型真实记忆能力的评估。为了解决这一问题,我们引入了OPIS,一个输入锚定的基准,严格将评估锚定到初始观察中的固定对象实例集,用于评估视频世界模型中的多对象记忆。OPIS数据集包含500个案例,涵盖真实世界、具身机器人和游戏世界领域,为12,672个刚性、铰接和可变形实例提供了密集的对象级标注。我们的对象中心评估器结合关联和显式可见性推理,基于对象运动学,分层测量对象(O)存在(P)、身份(I)和结构(S),并利用静态或动态评估轨道。在八个图像到视频或相机条件的世界模型中,我们提出的OPIS得分范围从48.65到56.01。随着参考清单从少于20个对象增长到超过40个对象,存在、身份和结构得分总体下降,平均身份得分从40.22降至23.11。结果表明,保留输入中的特定对象实例比生成合理的视觉元素要困难得多。
英文摘要:
Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded benchmark that strictly anchors the assessment to a fixed set of object instances from the initial observation for evaluating multi-object memory in video world models. The OPIS dataset comprises 500 cases across real-world, embodied-robotic, and game-world domains, providing dense object-level annotations for 12,672 rigid, articulated, and deformable instances. Our object-centric evaluator combines association and explicit visibility reasoning to hierarchically measure Object (O) Presence (P), Identity (I), and Structure (S), utilizing static or dynamic evaluation tracks based on object kinematics. Across eight image-to-video or camera-conditioned world models, our proposed OPIS scores range from 48.65 to 56.01. As the reference inventory grows from less than 20 to more than 40 objects, the Presence, Identity, and Structure scores show an overall decline, with the average Identity score falling from 40.22 to 23.11. The results demonstrate that preserving the particular object instances in the input is considerably harder than generating plausible visual elements.