arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30584cs.CV

基于合成课程学习的组合式时空视频定位方法

Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum

  • Monash University(莫纳什大学)
  • Southeast University(东南大学)
  • Shanghai University of Electric Power(上海电力大学)

机构由 AI 辅助整理,请以论文原文为准。

Xingjian Wang, Shijian Wang, Yibo Wang, Zihao Yu, Runhao Fu, Xuelian Cheng, Zongyuan Ge

AI总结:

本文针对现有时空视频定位模型忽略组合式查询的问题,提出CompSTVG任务,构建合成数据引擎与STVG-CompBench基准,引入CurrSTVG框架提升模型在组合式查询上的性能。

AI中文摘要:

尽管近期多模态大语言模型(MLLMs)在时空视频定位(STVG)领域取得了显著进展,但现有评估与训练数据主要聚焦于简单查询,很大程度上忽略了现实场景中普遍存在的组合式查询——这类查询需要通过联合推理目标的属性及其与其他实体的关系来区分目标。为填补这一空白,本文提出组合式时空视频定位(CompSTVG)任务,要求模型处理复杂文本查询,其中每个交织的属性和关系线索对区分目标至关重要。为规模化推进该任务,本文构建了一个合成数据引擎,该引擎将时空场景图作为难度衡量指标,将难度可控的查询合成转化为约束规划问题,生成用于评估和训练的难度分级数据。基于该引擎,本文推出STVG-CompBench基准,该基准按明确的难度层级分层,同时涵盖时间复杂度与空间干扰。在STVG-CompBench上对11种代表性STVG模型进行评估后发现,当前模型在组合式查询上表现不佳,出现了整体数据集级平均所掩盖的性能骤降现象。本文进一步构建合成训练数据,并提出课程强化学习框架CurrSTVG,该框架可带来持续提升,且在最具挑战性的组合式查询上观察到的改进幅度最大。

英文摘要:

Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on simple queries. They largely overlook the compositional queries prevalent in real-world scenarios, where a target must be disambiguated by jointly reasoning about its attributes and relations to other entities. To bridge this gap, we propose Compositional Spatio-Temporal Video Grounding (CompSTVG), a task that requires models to process complex textual queries where every intertwined attribute and relational cue is essential for disambiguation. To facilitate this task at scale, we build a synthetic data engine that leverages a spatio-temporal scene graph as a difficulty measure and casts difficulty-controlled query synthesis as a constraint programming problem, producing difficulty-graded data for both evaluation and training. Built on this engine, we introduce STVG-CompBench, a benchmark stratified by explicit difficulty levels that jointly capture temporal complexity and spatial interference. Evaluating 11 representative STVG models on STVG-CompBench reveals that current models perform poorly on compositional queries, exhibiting a sharp performance drop that is typically obscured by overall dataset-level averages. We further construct synthetic training data and propose CurrSTVG, a curriculum reinforcement learning framework that delivers consistent gains, with the largest improvements observed on the most challenging compositional queries.

补充信息

↑