arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17422cs.CV

TF-CADE:面向零样本时序动作检测的聚焦前景的文本-视频对齐

TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

Yearang Lee, Ho-Joong Kim, Seong-Whan Lee

首次发表
浏览论文内容

中文总结 AI 辅助

针对零样本时序动作检测中现有方法难以捕捉动作类别语义差异的问题,提出TF-CADE模型,通过ACA和CCR策略提升文本-视频对齐效果,实现分布内最优性能与跨数据集泛化能力。

中文摘要 AI 辅助

零样本时序动作检测(ZSTAD)旨在从未修剪视频中定位和识别未见过动作类别的动作实例。尽管现有方法通过改进架构级文本-视频对齐已展现出有效性,但仍难以捕捉动作类别间的语义差异,导致与文本无关的预测。为解决该问题,我们提出用于零样本时序动作检测器的文本-前景聚焦对齐(TF-CADE),其将文本信息与动作相关前景区域显式对齐。具体而言,我们引入动作聚焦聚合(ACA),该模块提取动作聚焦分数,将时间上有信息的视频片段聚合为前景加权视频嵌入。这种聚焦前景的对齐增强了文本与视频特征间的语义一致性,提升了类别间的判别性。此外,基于确定性的置信度重加权(CCR)策略,通过利用前景感知相似度优化每个片段的置信度分数,在推理时有效抑制无关动作类别。大量评估表明,我们的TF-CADE不仅在分布内设置下达到了最先进性能,还在针对未见过动作类别的跨数据集泛化中表现出色。

英文摘要

Zero-Shot Temporal Action Detection (ZSTAD) aims to lo- calize and recognize action instances from unseen action categories in untrimmed videos. Although existing meth- ods have shown effectiveness by advancing architectural text-video alignment, they still struggle with capturing se- mantic distinctions between action classes, resulting in text- irrelevant predictions. To address this issue, we propose a Text-Foreground Concentrated Alignment for zero-shot temporal action DEtector (TF-CADE) that explicitly aligns textual information with action-relevant foreground regions. Specifically, we introduce Action Concentrate Aggregation (ACA), which extracts action concentrate scores to aggregate temporally informative video segments into a foreground- weighted video embedding. This foreground concentrated alignment enhances the semantic consistency between text and video features and improves inter-class discriminabil- ity. In addition, a Certainty-based Confidence Re-weighting (CCR) strategy refines per-snippet confidence scores by lever- aging foreground-aware similarity, effectively suppressing irrelevant action classes during inference. Extensive evalua- tions show that our TF-CADE not only achieves state-of-the- art performance under in-distribution settings but also excels in cross-dataset generalization to unseen action classes.

发表机构

  • Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

↑