arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SV-TAD:用于高效时序动作检测的原生稀疏卷积

SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection

Ricardo Pizarro, Roberto Valle, José M. Buenaposada, Luis M. Bergasa, Luis Baumela

arXiv 2610.11579首次发表:更新:

发表机构

Universidad de Alcalá; Universidad Politécnica de Madrid; Universidad Rey Juan Carlos(阿尔卡拉大学; 马德里理工大学; 胡安·卡洛斯国王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出原生稀疏2D卷积,构建SV-TAD适配器框架,在降低VideoMAEv2-L计算量、提升推理速度的同时保持最优准确率,扩展后在InternVideoNext-L上也超越SOTA,还提升了ATTACH的细粒度装配检测性能。

AI 中文摘要

为了让拥有数十亿参数的视觉Transformer适配长视频理解,现有方法通常冻结主干网络并训练轻量卷积模块。这类方法虽能实现参数高效的训练,但无法降低推理阶段的计算量,在视频长度可扩展性方面仍存在较大问题。令牌选择可通过剪枝冗余令牌降低注意力成本,但会破坏卷积适配器所需的空间网格结构,这就需要进行代价高昂的密集重建,抵消了大部分潜在的加速效果。针对该问题,本文提出原生稀疏2D卷积,该原语首次让这些适配器能直接在动态剪枝后的令牌集上高效运行。我们将该原语集成到时序动作检测适配器框架SV-TAD中,使VideoMAEv2-L的计算量最多降低64%,推理速度提升2.2倍,同时在THUMOS-14和ActivityNet-1.3数据集上保持了当前最优的准确率。当扩展到InternVideoNext-L模型时,我们的方法以约一半的计算成本超越了之前的最优水平。此外,该稀疏形式天然支持辅助任务令牌,可提升ATTACH数据集上的细粒度装配检测性能。

英文摘要

To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules. While effective for parameter-efficient training, existing adapters do not reduce inference-time computation, leaving scalability with respect to video length largely unaddressed. Token selection can reduce attention cost by pruning redundant tokens, but it breaks the spatial grid structure required by convolutional adapters. This forces an expensive dense reconstruction, nullifying much of the potential speedup. We address this by introducing native sparse 2D convolutions, a primitive that allows these adapters, for the first time, to operate directly and efficiently on dynamically pruned token sets. We integrate this primitive into SV-TAD, an adapter framework for temporal action detection, reducing VideoMAEv2-L computation by up to 64% and achieving 2.2x faster inference, while maintaining state-of-the-art accuracy on THUMOS-14 and ActivityNet-1.3. When scaled to InternVideoNext-L, our approach surpasses the previous state of the art at roughly half its computational cost. Moreover, the sparse formulation naturally supports auxiliary task tokens, which improves fine-grained assembly detection on ATTACH.

Journal refECCV 2026. Lecture Notes in Computer Science, vol 17044. Springer

DOI:10.1007/978-3-032-37577-3_30

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑