发表机构
DreamX Team, Alibaba Group; Tongji University; Shanghai Jiao Tong University(阿里巴巴集团DreamX团队; 同济大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出R2M-Bench基准,通过同滚动过程的相对校准评估视频世界模型的重访记忆,其NMR与人类判断相关性良好,可减少慢动作捷径,DreamX-World-Memo表现最优。
AI 中文摘要
首次访问帧与返回帧之间的高相似度并不一定表明视频世界模型记住了该场景;中间的滚动过程可能只是几乎没有变化。这种模糊性使得绝对重访评分对渲染稳定性、重复内容和失败的运动很敏感。我们推出R2M-Bench(相对重访记忆基准,Relative Revisit Memory Benchmark),这是一个可观测的重访选择性一致性基准。对于每个检测到的返回,R2M-Bench将重访对与来自同一滚动过程的两个对照组进行比较:一个是匹配时间间隔的非重访对,用于测量通用时间稳定性;另一个是短程对,用于估计短程一致性。这些比较产生了MemoryGain(MG,即重访相对于时间基线的优势)和Normalized Memory Ratio(NMR,该指标将此优势按短程到基线的动态范围进行归一化)。R2M-Bench结合了100个参考场景与三条离开-返回轨迹,形成300个实例,并评估外观保真度、场景与对象身份、局部几何和持久状态。在七个基于动作的视频世界模型中,总体NMR与人类一致性判断的斯皮尔曼相关系数ρ=0.547(95%置信区间[0.45,0.63])。其与生成运动的模型内相关幅度为0.072,而原始重访相似度的该值为0.207,表明相对校准大幅减少了慢动作捷径。DreamX-World-Memo在评估的视频模型中实现了最高的总体NMR。这些结果共同支持同滚动过程相对校准作为区分重访特定一致性与通用时间稳定性的实用方法。
英文摘要
High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little. This ambiguity makes absolute revisit scores sensitive to rendering stability, repetitive content, and failed motion. We introduce \emph{R2M-Bench} (\textbf{R}elative \textbf{R}evisit \textbf{M}emory Benchmark), a benchmark of observable revisit-selective consistency. For every detected return, R2M-Bench compares the revisit pair with two controls from the same rollout: a gap-matched non-revisit pair that measures generic temporal stability and a short-range pair that estimates short-horizon consistency. These comparisons produce \emph{MemoryGain} (MG), the revisit advantage over the temporal baseline, and the \emph{Normalized Memory Ratio} (NMR), which normalizes this advantage by the short-to-baseline dynamic range. R2M-Bench combines 100 reference scenes with three leave-and-return trajectories to form 300 instances and evaluates appearance fidelity, scene and object identity, local geometry, and persistent state. Across seven action-conditioned video world models, Overall NMR correlates with human consistency judgments at Spearman's $ρ=0.547$ (95\% CI $[0.45,0.63]$). Its within-model correlation magnitude with generated motion is $0.072$, compared with $0.207$ for raw revisit similarity, indicating that relative calibration substantially reduces the slow-motion shortcut. DreamX-World-Memo achieves the highest Overall NMR among the evaluated video models. Together, these results support same-rollout relative calibration as a practical way to distinguish revisit-specific consistency from generic temporal stability.
CommentsCode: https://github.com/AMAP-ML/R2MBench