分组选择性状态空间滤波用于准确且流式的动作边界检测
Groupwise Selective State-Space Filtering for Accurate and Streaming Action Boundary Detection
浏览论文内容
中文总结 AI 辅助
本文提出一种基于分组选择性状态空间滤波的边界检测适配器,在预提取视频特征上学习时间表示,实现类别无关的流式动作边界检测,并在三个基准上取得良好F1分数。
中文摘要 AI 辅助
动作边界检测将未修剪视频划分为多个区间,而无需分配动作类别。我们提出了一种作用于预提取视频特征的边界检测适配器,通过分组选择性扫描学习时间表示。学习到的分组融合和时间建模将这些表示转换为转换分数,再解码为边界时间戳。该模型在边界时间监督下训练,是类别无关的,并在Breakfast、GTEA和50Salads数据集上使用时间容差和二分匹配进行评估,取得了0.457、0.622和0.611的边界F1分数。一种有状态变体支持特征流式推理,具有零神经前瞻、单样本峰值确认和有界内存。下游系统随后可以分配动作类别。
英文摘要
Action boundary detection partitions untrimmed video into intervals without assigning action classes. We present a boundary-detection adapter operating on pre-extracted video features, learning temporal representations via groupwise selective scans. Learned group fusion and temporal modeling convert these into transition scores, which are decoded into boundary timestamps. Trained with boundary-time supervision, the class-agnostic model is evaluated on Breakfast, GTEA, and 50Salads using temporal tolerances and bipartite matching, achieving boundary $F_1$ scores of 0.457, 0.622, and 0.611. A stateful variant enables feature-streaming inference with zero neural look-ahead, one-sample peak confirmation, and bounded memory. Downstream systems can subsequently assign s
发表机构
- Ankara Medipol University(安卡拉梅迪波尔大学)
机构由 AI 辅助整理,请以论文原文为准。