arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2610.04506cs.CVcs.AI

EgoExo-Next:视觉选项下一状态与跨视角推理的视觉语言模型基准

EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning

  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

Yutong Li, Molin Wang, Xiaotong Li, Yanyan Fang, Daoguo Dong, Ziyi Ye

AI总结:

本文提出EgoExo-Next基准,通过2,503个四选一问题评估视觉语言模型在动态视觉状态推理上的能力,发现最佳模型准确率仅43.81%,远低于人类的98.55%,凸显了现有模型在组合时间与跨视角推理上的显著局限。

AI中文摘要:

视觉语言模型(VLMs)越来越多地被用于第一人称视角和跨视角视频推理的评估,然而现有基准大多侧重于语义事件理解、时间关系或已观察视角之间的对应关系,而对其直接推理未来视觉状态的能力探索不足。我们引入了EgoExo-Next,一个用于动态视觉状态推理的视觉选项基准,其中模型必须识别观察到的动作轨迹随后如何呈现,而不仅仅是预测动作标签或文本描述。EgoExo-Next包含来自六个公开的第一人称和第一人称-第三人称视频来源的2,503个人工筛选的四选一问题,并包含四个相互关联的子任务,分别评估第一人称下一状态预测、双向第一人称-第三人称状态对应、第三人称下一状态预测,以及它们在从第一人称到第三人称下一状态任务中的组合。对专有、开源和空间推理视觉语言模型的广泛评估揭示了人类与模型之间的显著差距,最佳模型平均准确率为43.81%,而人类为98.55%,最大的性能下降发生在组合的第一人称到第三人称任务上。这些结果表明,当前的视觉语言模型在动态视觉状态推理方面仍然存在显著局限,尤其是在需要组合时间推进和跨视角推理时。该基准公开可用,网址为\url{this https URL}。

英文摘要:

Vision-language models (VLMs) are increasingly evaluated for egocentric and cross-view video reasoning, yet existing benchmarks largely focus on semantic event understanding, temporal relations, or correspondence between already observed views, leaving their ability to reason directly about future visual states underexplored. We introduce EgoExo-Next, a visual-option benchmark for dynamic visual-state reasoning, where models must identify how an observed action trajectory subsequently appears rather than predict only an action label or textual description. EgoExo-Next contains 2,503 human-curated four-choice questions from six public egocentric and ego--exo video sources and comprises four interconnected subtasks that evaluate egocentric next-state prediction, bidirectional ego--exo state correspondence, exocentric next-state prediction, and their composition in Ego-to-Exo Next-State. Extensive evaluation of proprietary, open-source, and spatial reasoning VLMs reveals a substantial human--model gap, with the best model achieving 43.81\% average accuracy compared with 98.55\% for humans, and the largest degradation occurring on the composed Ego-to-Exo task. These results suggest that current VLMs remain substantially limited in dynamic visual-state reasoning, particularly when temporal progression and cross-view reasoning must be composed. The benchmark is publicly available at \url{https://huggingface.co/datasets/yutongli2024/EgoExo-Next}.

↑