arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15555cs.CVcs.LG

RigidBench:评估视频生成模型中的刚体物理效果

RigidBench: Evaluating Rigid-Body Physics in Video Generation Models

  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

Swarnim Jain, Shangzhe Wu

中文总结 AI 辅助

该研究提出基于模拟器的RigidBench基准,评估视频生成模型的刚体物理效果,分析8个模型的表现,用5000个训练视频微调Wan 2.2 TI2V-5B并揭示其表征物体位置的机制。

中文摘要 AI 辅助

视频模型越来越多地被用于预测场景中接下来会发生什么,但常用于比较其输出的指标几乎无法判断预测物体的运动是否正确。运动、几何、身份、背景稳定性和视觉相似性可能会独立失效,而全帧评分通常会将这些错误混合在一起。我们推出 RigidBench,这是一个基于模拟器的基准,它将生成的视频续段与来自同一初始帧和运动描述的参考滚动结果进行比较。它的五项刚体任务涵盖不同的物体、材质、视角以及室内和室外场景,提供每帧掩码、深度、6自由度轨迹和接触信息用于评分。我们在相同的100个示例上用10项测量指标评估了8个模型,这些指标可将上述各个方面分开考量。得出的排名高度依赖所测量的内容:没有任何模型在全部10项指标上领先,且在模型均值中,更高的SSIM伴随着更大的3D轨迹误差(相关系数r=0.89)。RigidBench还包含5000个带有精确模拟器状态的训练视频,我们用这些视频对Wan 2.2 TI2V-5B进行微调与分析。完全微调可将3D轨迹误差降低约20%,同时几乎不改变SSIM;而教师强制探测和针对性干预显示,物体位置在Wan的扩散Transformer中被表征,并用于其去噪计算。

英文摘要

Video models are increasingly used to predict what happens next in a scene, yet the metrics commonly used to compare their outputs say little about whether the predicted objects move correctly. Motion, geometry, identity, background stability, and visual similarity can fail independently, but whole-frame scores often mix these errors together. We introduce RigidBench, a simulator-grounded benchmark that compares a generated continuation with a reference rollout from the same initial frame and motion description. Its five rigid-body tasks vary objects, materials, viewpoints, and indoor and outdoor scenes, with per-frame masks, depth, 6-DoF trajectories, and contacts available for scoring. We evaluate eight models on the same 100 examples with ten measurements that keep these aspects separate. The resulting rankings depend strongly on what is measured: no model leads on all ten, and across model means, higher SSIM accompanies larger 3D trajectory error (r = 0.89). RigidBench also includes 5,000 training videos with exact simulator state, which we use to fine-tune and analyze Wan 2.2 TI2V-5B. Full fine-tuning reduces 3D trajectory error by about 20% with almost no change in SSIM, while teacher-forced probes and targeted interventions show that object position is represented throughout Wan's diffusion transformer and used by its denoising computation.

补充信息

相关深度报道

↑