arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28666cs.CV

通过结构化视频提示提升视频-语言模型的时空推理能力

Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting

Sadegh Mohammadian

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出无需训练的结构化视频提示方法,在推理时通过为视频添加时空结构锚点,提升视频-语言模型在时空推理任务上的性能,为改进视频理解提供了实用方向。

中文摘要 AI 辅助

视频-语言模型(VLMs)在需要跟踪时间上的事件及将答案锚定到特定空间区域的任务上仍表现脆弱。我们认为该局限的部分原因可通过推理时更好地组织视觉证据来解决。我们提出结构化视频提示,这是一种无需训练的推理时方法,通过轻量空间结构和时间结构增强输入视频,为跨时空组织证据提供显式锚点,且无需更改模型权重、解码方式,也不改变主对比中的问题提示。我们在两个互补的视频基准和两个开源视频-语言模型上评估该方法。在这些设置中,结构化输入在若干情况下提升了性能,增益因模型和任务而异。我们的发现表明,VLMs的部分失败不仅源于推理能力,还源于推理时呈现视频证据的方式。这些结果凸显结构化视频提示是提升视频理解的简单实用方向。

英文摘要

Video-language models (VLMs) remain brittle on tasks that require tracking events over time and grounding answers in specific spatial regions. We propose that part of this limitation can be addressed through better organization of visual evidence at inference time. We introduce structured video prompting, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding and without altering the question prompt in the main comparison. We evaluate this approach on two complementary video benchmarks and two open video-language models. Across these settings, structured inputs improve performance in several cases, with gains varying by model and task. Our findings suggest that some failures of VLMs arise not only from reasoning capacity, but also from how video evidence is presented at inference time. These results highlight structured video prompting as a simple and practical direction for improving video understanding.

发表机构

  • Sharif University of Technology(谢里夫理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑