发表机构
Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究时空视频定位问题,提出ScanFocus从粗到细框架,利用统一融合编码器和轻量级模块生成粗略提议,再用语义引导时间聚合器恢复细节,实验表明该方法性能优于以往方法。
AI 中文摘要
时空视频定位(STVG)旨在从视频流中检索由自然语言表达描述的特定对象的视觉轨迹。然而,大多数先进方法难以在全局上下文建模和精确边界定位之间取得平衡。由于处理长视频的计算成本过高,这些方法通常采用低速率时间下采样和隐式运动建模,这不可避免地抑制了高频边界线索并忽略了精确边界描绘所需的显式帧间依赖性。为了解决这些限制,我们提出了ScanFocus,一种新颖的从粗到细框架,将STVG任务解耦为全局时空扫描和局部边界聚焦。具体来说,我们使用统一的视觉语言融合编码器结合轻量级可变形语义运动融合模块来有效地对齐多模态特征并生成粗略提议。为了恢复被抑制的细粒度细节,我们在细化阶段引入了语义引导时间聚合器(SGTA)。通过在粗边界周围密集采样,SGTA在语义引导下显式建模短期时间交互,捕获快速运动变化以进行精确时间戳回归。在三个广泛使用的基准上进行的大量实验证明了我们提出的方法优于以前的方法。代码将在这个https URL上发布。
英文摘要
Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression. However, most advanced methods struggle to balance global context modeling with precise boundary localization. Due to the prohibitive computational costs of processing long videos, these approaches typically resort to low-rate temporal downsampling and implicit motion modeling. This inevitably suppresses high-frequency boundary cues and neglects the explicit inter-frame dependencies required for precise boundary delineation. To address these limitations, we present \textbf{ScanFocus}, a novel coarse-to-fine framework that decouples the STVG task into a global spatio-temporal scan and a local boundary focus. Specifically, we utilize a unified vision-language fusion encoder combined with a lightweight Deformable Semantic-Motion Fusion module to efficiently align multimodal features and generate coarse proposals. To recover the suppressed fine-grained details, we introduce the Semantic-Guided Temporal Aggregator (SGTA) in the refinement stage. By densely sampling around coarse boundaries, SGTA explicitly models short-term temporal interactions under semantic guidance, capturing rapid motion changes for precise timestamp regression. Extensive experiments on three widely used benchmarks demonstrate the performance superiority of our proposed method over previous approaches. Code will be released at https://github.com/TenMinutes209/ScanFocus.
Commentsthis paper has already been accepted by ECCV 2026