S2AFormer: Strip Self-Attention for Efficient Vision Transformer
S2AFormer:用于高效视觉变换器的条带自注意力
机构 * Faculty of Engineering and Information Technology, University of Technology Sydney(工程与信息技术学院,悉尼大学) ; Bionic Vision System Laboratory, State Key Laboratory of Transducer Technology, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences(生物视觉系统实验室,传感技术国家重点实验室,上海微系统与信息技术研究所,中国科学院) ; PCA Lab, Key Lab of Intelligent Perception and Systems for High-Dimensional Information of Ministry of Education, School of Computer Science and Engineering, Nanjing University of Science and Technology(PCA实验室,教育部高维信息智能感知与系统重点实验室,南京理工大学计算机科学与工程学院) ; Research Center for Industries of the Future and the School of Engineering, Westlake University(未来产业研究中心和工程学院,西湖大学) ; OPPO Research, Seattle, WA 98101 USA(OPPO研究,美国西雅图华盛顿州98101)
AI总结 S2AFormer通过引入条带自注意力机制,有效整合CNN的局部感知与Transformer的全局上下文建模,提升视觉变换器的效率与准确性。
Comments Accepted by IEEE-TIP, 14 pages, 8 figures, 9 tables