arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VideoSEMA:用于视频理解的一种可扩展且高效的类曼巴注意力机制

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

Nhat Thanh Tran, Fanghui Xue, Shuai Zhang, Jiancheng Lyu, Yunling Zheng, Yingyong Qi, Jack Xin

arXiv 2607.14711首次发表:更新:

发表机构

University of California, Irvine; Qualcomm AI Research; Qualcomm Technologies, Inc.(加利福尼亚大学欧文分校; 高通人工智能研究中心; 高通技术公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视频理解提出VideoSEMA模型,核心方法是采用时空注意力分割,包含空间类曼巴注意力模块和时间注意力。主要贡献是在多数据集上表现优异,参数规模相当时准确率领先,分辨率提升时退化更平缓,扩展到长视频也有前景。

AI 中文摘要

我们提出了用于视频理解(分类)的时空注意力分割模型VideoSEMA,它由空间上可扩展且高效的类曼巴注意力(SEMA)模块和时间上的softmax时间注意力组成。在每一帧中,SEMA注意力在曼巴宏架构中并行应用局部窗口注意力和全局平均,即类曼巴。在一定秩条件下,证明了计算成本更低的时空注意力分割等同于全时空注意力。在基准K400数据集上,VideoSEMA优于更复杂的视觉Transformer和曼巴模型。在基准SSv2数据上,VideoSEMA在相似参数规模模型中top-1准确率领先。在K400上图像分辨率从标准的$224^2$提升到$1024^2$且未微调时,VideoSEMA在准确率上比VideoMamba退化更平缓。将VideoSEMA扩展到更长视频并采用扩张/稀疏时间注意力很有前景。

英文摘要

We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local window attention in parallel with a global averaging in a Mamba macro-architecture, which is called Mamba-like. Under certain rank conditions, we prove that the computationally cheaper split space-time attention is equivalent to full space-time attention. On benchmark K400 data sets, VideoSEMA out-performs heavier vision transformer and Mamba models. On benchmark SSv2 data, VideoSEMA leads in top-1 accuracy among models of similar parameter sizes. As image resolution scales up from standard $224^2$ to $1024^2$ on K400 and without fine-tuning, VideoSEMA degrades much more gracefully than VideoMamba in accuracy. It is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.

Comments15 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑