arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越视觉质量:基于EgoGenEval评估自我运动下的物理一致性

Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEval

Yilin Long, Chenming Zhu, Zitang Gou, Jingli Lin, Tai Wang

arXiv 2609.11172首次发表:更新:

发表机构

Shanghai AI Laboratory; Fudan University; The University of Hong Kong; Shanghai Jiao Tong University(上海人工智能实验室; 复旦大学; 香港大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉生成器在自我运动下物理一致性不足的问题,提出EgoGenEval基准和EgoGen-Train数据,发现成对监督难以同时提升相机运动接地与场景状态保持,需转向轨迹中心范式。

AI 中文摘要

最近的视觉生成器能产生高保真图像,但在自我运动下常常违反物理一致性,限制了它们在空间推理和具身规划中的应用。现有基准大多关注孤立图像或单步质量,使这一挑战未被充分探索。我们提出了EgoGenEval,一个基于几何、无需姿态的基准,旨在评估视觉生成器在自我运动下的物理一致性,并将研究分为两部分。(1)EgoGenEval包含1,400个案例和2,360个目标视图,涵盖单步和多步自我运动。它分别测量相机运动接地(CMG)和场景状态保持(SSP),两个指标均与盲人判断进行验证。评估16个无需姿态的生成器以及两个以姿态为条件的参考模型,揭示当前模型难以在保持场景状态的同时执行相机运动,且没有系统能在两个轴上都表现良好。(2)为了检验基准数据是否能提升这些能力,我们从相同的几何接地流程构建了EgoGen-Train,并进行了受控的SFT研究。这些研究表明,成对监督不能可靠地同时改善相机运动接地和场景状态保持:即使在完整训练池和最长的预算下,场景保持的提升也只是相机运动提升的一小部分。这表明成对教师强制目标本身是约束瓶颈,促使我们转向一种以轨迹为中心的范式,将自条件滚动与显式姿态和可见性监督相结合。

英文摘要

Recent visual generators produce high-fidelity images yet often violate physical consistency under ego-motion, limiting their use for spatial reasoning and embodied planning. Existing benchmarks largely focus on isolated images or single-step quality, leaving this challenge underexplored. We introduce EgoGenEval, a geometry-grounded, pose-free benchmark designed to evaluate the physical consistency of visual generators under ego-motion, and organize our study into two parts. (1) EgoGenEval contains 1,400 cases and 2,360 target views spanning single-step and multi-step ego-motion. It separately measures Camera Motion Grounding (CMG) and Scene State Preservation (SSP), with both metrics validated against blinded human judgments. Evaluating 16 pose-free generators together with two pose-conditioned references reveals that current models struggle to execute camera motion while maintaining scene state, and that no system performs well on both axes at once. (2) To examine whether benchmark-derived data can improve these capabilities, we build EgoGen-Train from the same geometry-grounded pipeline and run controlled SFT studies. These show that pairwise supervision does not reliably improve camera-motion grounding and scene-state preservation together: even at the full training pool and the longest budget, scene preservation gains a fraction of what camera motion does. This points to the pairwise teacher-forced objective itself as the binding constraint, motivating a trajectory-centric paradigm that couples self-conditioned rollouts with explicit pose and visibility supervision.

Comments39 pages, 10 figures, and 18 tables. Code: https://github.com/InternRobotics/EgoGenEval

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑