arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27929cs.CV

用于通用视频理解的无训练时间抽象

Training-Free Temporal Abstraction for General Video Understanding

Etienne Casanova, Sevan Brodjian, Pietro Perona

首次发表
浏览论文内容

中文总结 AI 辅助

提出无训练的STITCH方法,将视频划分为语义时间块,在三个视频理解任务上表现具竞争力,为通用视频理解提供新方向。

中文摘要 AI 辅助

逐帧分析视频的成本很高,但许多视频理解任务依赖于识别相关时刻的位置。系统可能需要找到动作发生变化的时间、定位句子所描述的片段,或为视觉语言模型(VLM)选择少量帧。现有方法通常通过特定任务的训练数据或专用架构分别解决这些问题。本研究探讨预训练的视频文本模型是否能提供足够的时间结构,以同时支持多个此类任务。我们提出STITCH,一种无训练方法,它将视频划分为具有语义意义的时间块。STITCH使用冻结的视频文本骨干网络嵌入短视频窗口,并检测所得嵌入序列中的变化。这些块每个视频仅计算一次,可跨任务重复使用。我们在通用事件边界检测、基于语言的时刻检索以及长视频VLM推理的帧选择这三个场景中评估STITCH。在所有三个场景中,STITCH与更专用的方法相比仍具有竞争力,且无需特定任务的训练;当只能处理少量帧或标记时,STITCH的优势尤为明显。这些结果表明,可复用的时间抽象是通用视频理解的一个有前景的方向,它能将密集的视频流一次性转换为语义单元,供下游系统进行定位、检索、采样或推理。

英文摘要

Videos are expensive to analyze frame by frame, yet many video understanding tasks depend on knowing where relevant moments occur. A system may need to find when an action changes, locate the segment described by a sentence, or choose a few frames for a vision-language model. Existing methods often solve these problems separately, using task-specific training data or specialized architectures. We study whether a pretrained video-text model can provide enough temporal structure to support several of these tasks at once. We present STITCH, a training-free method that divides a video into semantically meaningful temporal chunks. STITCH embeds short video windows with a frozen video-text backbone and detects changes in the resulting embedding sequence. These chunks are computed once per video and reused across tasks. We evaluate STITCH on generic event boundary detection, language-based moment retrieval, and frame selection for long-video VLM reasoning. Across all three settings, STITCH remains competitive with more specialized methods while requiring no task-specific training, with especially clear gains when only a small number of frames or tokens can be processed. These results suggest that reusable temporal abstraction is a promising direction for general video understanding, allowing dense video streams to be converted once into semantic units that can be localized, retrieved, sampled, or reasoned over by downstream systems.

发表机构

  • California Institute of Technology(加州理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑