arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WorldExam:从表观外观到内在反应性对世界模型进行基准测试

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Yuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang

arXiv 2608.02603首次发表:更新:

发表机构

CASIA; SLAI; CUHK; AMAP; THU(中国科学院自动化研究所; 智能科学与技术实验室; 香港中文大学; 中国农业科学院; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出WorldExam基准,评估20个模型在视觉质量等四层级的表现,发现不同驱动模型能力有明显分化,无模型兼具广泛任务覆盖与稳定性能,高视觉质量不代表内在反应性强。

AI 中文摘要

可控视频生成模型正日益被开发为世界模型,因此对其作为世界模型的评估已超出生成视频的表观外观,延伸至其描绘世界的内在反应性:即从场景状态推断世界应如何反应并生成输入未明确描述的合理结果的能力。然而现有基准主要通过检查请求的动作和交互结果是否实现来评估视觉质量或明确指令完成情况,却未充分考察内在反应性。我们推出WorldExam,这是一个涵盖四个层级的分层诊断基准:视觉质量、控制依从性、空间一致性和世界反应性。它包含8个专用任务的1474个案例,支持对相机驱动、动作驱动和语言驱动模型范式的统一评估。世界反应性层级评估输入中未明确指定的场景条件反应和目标导向行为。对20个代表性模型的评估显示出明显的能力分化:相机驱动模型擅长相机控制,但其接口不支持动态交互;动作驱动模型能更精准地控制主体,但常使世界处于无反应状态;语言驱动模型在交互上表现更好,但对复杂控制的依从性较差。没有模型同时具备广泛的任务覆盖范围和始终优异的性能,表明高视觉质量和明确指令完成情况并不能保证内在反应性。

英文摘要

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.

CommentsProject Website: https://WorldExam.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑