arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06691cs.CV

CoDAT:面向高效边缘动作识别的、具备低成本时序建模的协同双注意力Transformer

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

发表机构国立虎尾科技大学 · 马朗穆罕默迪亚大学 · 台湾科技大学
另 1 家 · 查看机构详情
  • National Formosa University(国立虎尾科技大学)
  • University of Muhammadiyah Malang(马朗穆罕默迪亚大学)
  • National Taiwan University of Science and Technology(台湾科技大学)
  • National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu, Wen-Kai Kuo, Jing-Ming Guo, Jun-Wei Hsieh

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出CoDAT协同双注意力Transformer,通过轻量双分支模块与TShift实现低成本时序建模,在Jetson AGX Orin等边缘设备上,兼顾图像与动作识别的能效-准确率平衡,运行速度远超同类模型。

中文摘要 AI 辅助

物联网(IoT)边缘设备上的实时人体动作识别,要求模型在严格的延迟、内存和功耗约束内捕捉丰富的时空线索。当前的3D CNN、视频Transformer和基于移位的ViT虽能实现高准确率,但计算成本过高,无法部署在边缘IoT设备中。本文提出CoDAT,即协同双注意力Transformer,它用轻量双分支模块替代传统多头注意力:空间卷积注意力(SCA)用于局部聚合,跨步单头注意力(SSHA)用于全局上下文建模。SSHA通过基于跨步的稀疏投影,同时压缩查询、键和值张量的空间分辨率与通道维度,再以显著降低的成本融合生成的全局与局部特征。为实现跨帧的时序通信,每个模块中嵌入了无参数的TShift模块。在Jetson AGX Orin和Raspberry Pi 5上开展的大量实验表明,CoDAT在图像识别和动作识别中均实现了能效-准确率的平衡。在ImageNet-1K上,CoDAT-M在准确率相当的情况下,运行速度比EfficientViT384和FastViT-S12快2倍;CoDAT-L的参数比ViT-S少3倍,吞吐量却高2倍,性能与ViT-S相当。在Kinetics-400和MA-52上,CoDAT在达到与最先进的CNN、Transformer和混合基线相当的Top-1准确率的同时,运行速度比VSwin-T快达2.9倍、比ViT-Temporal-Shift变体快2倍、比UniFormer-B快5倍。在UCF-101上,CoDAT-S384的准确率与TokShift和LAPS相当,运行速度却快6倍,所需FLOPs最多减少13倍,为边缘IoT感知系统中的实时动作识别建立了能效-准确率的平衡。代码可在此https URL获取。

英文摘要

Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency, memory, and power envelopes. Current 3D CNNs, video transformers, and shift-based ViT deliver high accuracy but come at computational costs that preclude edge IoT deployment. This paper proposes CoDAT, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context. SSHA jointly compresses the spatial resolution and channel dimensions of the query, key, and value tensors via stride-based sparse projection, then fuses the resulting global and local features at a markedly reduced cost. To enable temporal communication across frames, a parameter-free TShift module is embedded in each block. Extensive experiments on Jetson AGX Orin and Raspberry Pi 5 demonstrate that CoDAT achieves an energy-accuracy balance in both image and action recognition. On ImageNet-1K, CoDAT-M runs 2x faster than EfficientViT384 and FastViT-S12 at comparable accuracy, and CoDAT-L matches ViT-S with 3x fewer parameters at 2x higher throughput. On Kinetics-400 and MA-52, CoDAT achieves competitive Top-1 accuracy against state-of-the-art CNN, transformer, and hybrid baselines while running up to 2.9x faster than VSwin-T, 2x faster than ViT-Temporal-Shift variants, and 5x faster than UniFormer-B. On UCF-101, CoDAT-S384 matches TokShift and LAPS while being 6x faster and requiring up to 13x fewer FLOPs, establishing an efficiency-accuracy balance for real-time action recognition in edge IoT perception systems. Code is available at https://github.com/novendrastywn/CoDAT .

补充信息

↑