arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Video-HolmesV2:多模态大语言模型能否在长视频中利用时空视听证据进行推理?

Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?

Zhaoyang Wei, Zipeng Wang, Yushe Cao, Chenhui Qiang, Shuaibing Cheng, Xuesong Yang, Sen Nie, Bowen Jiang, Wenchao Ding, Yanchao Hao, Zheng Wei, Xuehui Yu, Zhenjun Han

arXiv 2609.17248首次发表:更新:

发表机构

University of Chinese Academy of Sciences; Tencent; Tsinghua University(中国科学院大学; 腾讯; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视频推理中视觉中心化评估与上下文处理低效问题,提出Video-HolmesV2基准及音频文本引导的令牌压缩框架,实现基于时空视听证据的评估,提升跨模态推理性能。

AI 中文摘要

多模态大语言模型已展现出令人印象深刻的视频理解能力,然而其处理长篇叙事推理的能力常常被以视觉为中心的评估和低效的上下文处理所掩盖。现有基准过度依赖视觉启发式方法,同时边缘化听觉线索,实际上将模型降级为“沉默的观察者”,绕过了真正的跨模态推理。此外,标准密集采样造成了证据与上下文之间的权衡:增加帧数以捕获证据不可避免地导致注意力分散和令牌爆炸。为弥合这些差距,我们提出了Video-HolmesV2,一个专为深度视听耦合设计的新型基准。与以往工作不同,它强制执行基于证据的评估,要求模型用精确的时空视听证据来证明答案,从而减少猜测和幻觉证据的混淆影响。为此,我们引入了:(1)一个多模型交叉验证流程以确保任务严谨性;(2)一个时空证据感知指标用于细粒度校准。此外,我们提出了一个音频文本引导的令牌压缩框架。通过将任务意图与听觉锚点融合,我们的方法提炼出高价值的推理线索以缓解长上下文噪声。在我们的评估中,即使是强大的专有模型准确率也低于60%,而我们的方法优于可比较的开源全模态模型。

英文摘要

Multimodal Large Language Models have demonstrated impressive video understanding, yet their ability to reason over long-form narratives is often masked by visual-centric evaluations and inefficient context processing. Existing benchmarks over-rely on visual heuristics while marginalizing auditory cues, effectively reducing models to "silent observers" that bypass genuine cross-modal reasoning. Moreover, standard dense sampling creates an evidence-context trade-off: increasing frames to capture evidence inevitably leads to attention distraction and token explosion. To bridge these gaps, we present Video-HolmesV2, a novel benchmark designed for Deep Audio-Visual Coupling. Unlike previous works, it enforces an Evidence-Based Evaluation, requiring models to justify answers with precise spatio-temporal audio-visual evidence, thereby reducing confounding effects of guessing and hallucinated evidence. To support this, we introduce: (1) a Multi-Model Cross-Verification pipeline to ensure task rigor; (2) a Spatio-temporal Evidence-Aware Metric for fine-grained calibration. Furthermore, we propose an Audio-Text Guided Token Compression framework. By fusing task intent with auditory anchors, our method distills high-value reasoning cues to mitigate long-context noise. In our evaluation, even strong proprietary models achieve below 60% accuracy, while our approach outperforms comparable open-source omni-models.

CommentsAccepted by ECCV2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑