arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30952cs.CVcs.AI

MVVBench:视觉语言模型中的4D推理基准

MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim, Jaeho Choi, Byung-Hoon Kim

首次发表
浏览论文内容

中文总结 AI 辅助

MVVBench是一个多视角视频推理基准,通过设计单目模糊的问题,评估视觉语言模型跨视角和跨时间的4D推理能力,并探索推理时引导和强化学习等策略以提升性能。

中文摘要 AI 辅助

多视角视频理解需要整合来自多个、通常不重叠的摄像头流的空间和时间证据:跟踪实体在视点间的转换,跨时间对齐事件,并推理潜在的4D连续性,而非任何单个可见帧。我们引入了MVVBench,一个基于真实世界多摄像头数据集构建的多视角视频推理基准。问题被精心设计为在视角和时间轴上都具有单目模糊性:每个问题都无法从指定输入集中的任何单一视角回答,且大多数问题也无法从任何单一时刻回答。每个问题只有通过跨视角和跨时间的联合推理才能唯一解决。MVVBench涵盖多样的动态场景,并探测六种能力:隐式/显式属性识别、隐式/显式相对距离、相对相机姿态和组合计数,配有手工编写的问答和严格的验证。除了基准测试外,我们还提供了对当前视觉语言模型何时以及为何成功或失败的广泛分析,描述了由于时间定位错误、跨视角身份断裂和脆弱的多跳推理导致的错误。然后,我们研究了推理时的引导策略,以解锁潜在的多视角能力——特定任务的思维链支架和结构化的跨视角证据聚合——在无需重新训练的情况下产生了显著提升。最后,我们提供了初步证据表明,使用可验证奖励的强化学习可以在基础模型中激发一些潜在的多视角能力,指向训练时方法作为未来工作的有前景方向。总之,MVVBench提供了对4D多视角推理的严格评估,并为未来朝着可靠具身感知的进展奠定了基础。

英文摘要

Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be monocular-ambiguous along both the view and the temporal axis: each question is unanswerable from any single view in the designated input set, and the majority are further unanswerable from any single moment. Each question becomes uniquely solvable only by jointly reasoning across views and across time. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, with human-authored QA and rigorous verification. Beyond benchmarking, we provide an extensive analysis of when and why current vision language models succeed or fail, characterizing errors due to temporal mis-localization, cross-view identity breaks, and brittle multi-hop reasoning. We then study inference-time elicitation strategies that unlock latent multi-view competence---task-specific chain-of-thought scaffolds and structured cross-view evidence aggregation---yielding substantial gains without retraining. Finally, we present preliminary evidence that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work. Together, MVVBench offers a rigorous evaluation of 4D multi-view reasoning and a foundation for future progress toward reliable embodied perception.

发表机构

  • EverEx
  • Korea University(高丽大学)
  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑