arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33400cs.CV

分组选择性状态空间滤波用于准确且流式的动作边界检测

Groupwise Selective State-Space Filtering for Accurate and Streaming Action Boundary Detection

Mustafa Bora Çelik

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种基于分组选择性状态空间滤波的边界检测适配器,在预提取视频特征上学习时间表示,实现类别无关的流式动作边界检测,并在三个基准上取得良好F1分数。

中文摘要 AI 辅助

动作边界检测将未修剪视频划分为多个区间,而无需分配动作类别。我们提出了一种作用于预提取视频特征的边界检测适配器,通过分组选择性扫描学习时间表示。学习到的分组融合和时间建模将这些表示转换为转换分数,再解码为边界时间戳。该模型在边界时间监督下训练,是类别无关的,并在Breakfast、GTEA和50Salads数据集上使用时间容差和二分匹配进行评估,取得了0.457、0.622和0.611的边界F1分数。一种有状态变体支持特征流式推理,具有零神经前瞻、单样本峰值确认和有界内存。下游系统随后可以分配动作类别。

英文摘要

Action boundary detection partitions untrimmed video into intervals without assigning action classes. We present a boundary-detection adapter operating on pre-extracted video features, learning temporal representations via groupwise selective scans. Learned group fusion and temporal modeling convert these into transition scores, which are decoded into boundary timestamps. Trained with boundary-time supervision, the class-agnostic model is evaluated on Breakfast, GTEA, and 50Salads using temporal tolerances and bipartite matching, achieving boundary $F_1$ scores of 0.457, 0.622, and 0.611. A stateful variant enables feature-streaming inference with zero neural look-ahead, one-sample peak confirmation, and bounded memory. Downstream systems can subsequently assign s

发表机构

  • Ankara Medipol University(安卡拉梅迪波尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑