arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17112cs.CV

不是又一个文本基准:将“视觉”重新放回大型视频模型的视觉问答中

Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models

发表机构卡内基梅隆大学 · 德克萨斯A&M大学 · 约翰斯·霍普金斯大学
另 3 家 · 查看机构详情
  • Carnegie Mellon University(卡内基梅隆大学)
  • Texas A&M University(德克萨斯A&M大学)
  • Johns Hopkins University(约翰斯·霍普金斯大学)
  • University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
  • UiT The Arctic University of Norway(挪威北极圈大学)
  • University of Copenhagen(哥本哈根大学)

机构由 AI 辅助整理,请以论文原文为准。

Rwiddhi Chakraborty, Yinong, Wang, Cheng Zhang, Fan Bai, Zhuoran You, Michael Kampffmeyer, Yong Jae Lee, Fernando De la Torre, Robert Jenssen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出三个以视觉为中心的基准(时间帧检索、视频未来预测、因果记忆扭曲),发现大型视频模型在视觉查询推理上显著弱于文本推理,并给出改进方向。

中文摘要 AI 辅助

大型视频模型在广泛的视觉问答任务上展现了令人印象深刻的表现,这得益于强大的预训练文本和视觉编码器的兴起。这类模型的有用性也在广泛的基准上得到了证明,但有一个重要的注意事项——这些基准中的主导方法是通过文本选项来评估多项选择推理。这是测试这些模型中基于文本推理的自然方式,并已在社区中引发了关于模型行为的重要见解。在这项工作中,我们提出了一个不同的问题——当评估模态是视觉而非文本时,会发生什么?我们引入了三个新的以视觉为中心的评估基准,分别涉及时间帧检索、视频未来预测和因果记忆扭曲,所有这些都旨在评估大型视频模型中的视觉理解能力。我们的方法补充了现有评估前沿模型视频理解的方法。我们表明,当前的前沿模型在尝试通过视觉查询而非文本进行推理时表现出显著的弱点。我们以一个扩展的分析部分作为结论,为大型视频模型未来在视觉理解方面的改进提供了方向。

英文摘要

Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The usefulness of such models have also been demonstrated on a wide range of benchmarks, with an important caveat - the dominant approach in these benchmarks evaluates multiple choice reasoning via text options. This is a natural way to test text-based reasoning in these models, and has led to significant insights regarding model behavior in the community. In this work, we ask a different question - what happens when the evaluation modality is visual, rather than text? We introduce three new vision-centric evaluation benchmarks in temporal frame retrieval, video future prediction, and causal memory distortion, all designed around evaluating visual understanding capabilities in large video models. Our approach complements the existing approaches to evaluate video understanding in frontier models. We show that current frontier models exhibit significant weakness when attempting to reason through visual queries, rather than text. We conclude with an extended analysis section that provides pointers for future improvements in visual understanding for large video models.

↑