arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OvDSGG:端到端开放词汇动态场景图生成

OvDSGG: End-to-End Open-Vocabulary Dynamic Scene Graph Generation

John Helsby, Yi Yang, Bodo Rosenhahn, Michael Ying Yang

arXiv 2608.14835首次发表:更新:

发表机构

Meta; University of Bath; Leibniz Universität Hannover(Meta; 巴斯大学; 汉诺威莱布尼茨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OvDSGG是首个开放词汇DSGG的端到端框架,通过空间主干、时间主干及相关模块实现开放词汇识别,在开放词汇指标上显著优于基线,封闭集性能仍具竞争力,且构建了对应基准并公开代码。

AI 中文摘要

动态场景图(DSG)以〈主体、谓词、客体〉三元组形式捕获视频中的时空交互,为视频字幕、视频问答、动作分析等下游任务提供支撑。然而,端到端动态场景图生成(DSGG)方法属于封闭集,仅能识别固定训练词汇表中的客体和谓词,难以应对稀有概念的长尾分布,严重限制了其实际应用。现有开放词汇模型通常继承预训练大语言模型,导致多阶段训练与推理,成本高昂。我们提出OvDSGG,首个开放词汇DSGG的端到端框架。OvDSGG基于开放词汇空间主干和时间主干构建,进一步提出连接两者的三元组特征提取模块,以及视觉-语言对齐模块,该模块通过在联合视觉-语言特征空间中学习自适应决策边界来保持开放词汇识别能力,无需现有方法中昂贵的知识蒸馏。我们还基于Action Genome构建了严格的开放词汇DSGG基准,客体和谓词均采用不相交的基类/新类划分。OvDSGG在所有指标上均显著优于开放词汇基线,零样本Recall@K得分比次优基线高10.0至20.4个百分点,同时在封闭集DSGG上仍与最先进模型具有竞争力。代码和基准可在该URL公开获取。

英文摘要

Dynamic scene graphs (DSGs) capture spatio-temporal interactions across videos as $\langle$subject, predicate, object$\rangle$ triplets, and underpin downstream tasks such as video captioning, video question answering, and action analysis. However, end-to-end dynamic scene graph generation (DSGG) methods are closed-set: they recognize only objects and predicates from a fixed training vocabulary and struggle with the long-tailed distribution of rare concepts, severely limiting their real-world applicability. Existing open-vocabulary models typically inherit pretrained large language models, resulting in multi-stage training and inference with substantial cost. We introduce OvDSGG, the first end-to-end framework for open-vocabulary DSGG. OvDSGG builds on top of an open-vocabulary Spatial Backbone and a Temporal Backbone; we further propose a Triplet Feature Extraction Module that bridges them, and a Visual-Language Alignment Module that preserves open-vocabulary recognition by learning an adaptive decision boundary in the joint visual-language feature space, without expensive knowledge distillation in existing methods. We further introduce a rigorous open-vocabulary DSGG benchmark adapted from Action Genome, with disjoint Base/Novel splits for both objects and predicates. OvDSGG significantly outperforms open-vocabulary baselines across all metrics, with zero-shot Recall@$K$ scores 10.0--20.4 percentage point higher than the next-best baseline, while on closed-set DSGG remaining competitive with state-of-the-art models. Code and benchmark are publicly available at https://github.com/jhelsby/OvDSGG/.

CommentsECCVW'26 CONTEXTUS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑