发表机构
Xi’an Jiaotong University; The Hong Kong Polytechnic University; OPPO Research(西安交通大学; 香港理工大学; OPPO研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对等变网络计算效率低的问题,提出Flash EQ-Linear算法,结合傅里叶卷积定理等降低复杂度,提供CUDA内核,在算子和网络层面均实现加速,使等变网络在多方面首次严格优于非等变网络。
AI 中文摘要
等变网络通过权重共享将几何对称性作为结构先验进行嵌入,在视觉任务中实现了显著的参数效率。然而,这种参数效率并未转化为计算效率,现有实现将结构化权重展开为密集矩阵并分配给通用密集内核,导致等变层的浮点运算次数不小于非等变层。本文观察到等变线性(EQ-Linear)层本质上是沿组维度的循环卷积与沿通道维度的线性变换的组合。基于此,提出了Flash EQ-Linear,通过沿组维度结合傅里叶卷积定理和实离散傅里叶变换的共轭对称性,将复杂度从\(\mathcal{O}(NDC)\)降低到\(\mathcal{O}(NDC/T)\)。还提供了专用的CUDA内核,涵盖前向和反向传播以及FP32和FP16精度。在算子层面,Flash EQ-Linear比PyTorch的实现实现了高达2倍的前向加速;在网络层面,Flash EQ-ViT和Flash EQ-Swin比等变和非等变基线实现了高达1.7倍的端到端加速。据我们所知,这是等变网络首次在精度、参数效率和推理速度这三个轴上同时严格优于非等变网络。
英文摘要
Equivariant networks embed geometric symmetries as structural priors through weight sharing, achieving remarkable parameter efficiency across vision tasks. However, this parameter efficiency does not translate into compute efficiency: most existing implementations unroll the structured weights into dense matrices and dispatch them to generic dense kernels, so an equivariant layer costs no fewer MACs than its non-equivariant counterpart. In this paper, we observe that the equivariant linear (EQ-Linear) layer---the most fundamental and frequently used module in modern equivariant architectures---is essentially a circular convolution along the group dimension composed with a linear transform along the channel dimension. Building on this observation, we propose Flash EQ-Linear, an exact acceleration algorithm that reduces the cost to $2(T-1)/T^2$ of the original dense formulation ($T$ is the equivariant group size) by combining the Fourier convolution theorem along the group dimension with the conjugate symmetry of the real DFT. To translate these computational savings into wall-clock speedups, we further develop dedicated CUDA kernels for the $\mathrm{p}4$ group. At the operator level, Flash EQ-Linear achieves up to $2.1\times$ forward speedup over PyTorch's highly optimized F.linear; at the network level, Flash EQ-ViT achieves up to ${1.7\times}$ end-to-end speedup over both equivariant and non-equivariant baselines. As an operator-level acceleration algorithm, Flash EQ-Linear provides plug-and-play acceleration for diverse pretrained equivariant models, including EQ-ViT, EQ-Swin, EQ-VMamba, and EQ-INR, without retraining or architectural changes. Code is available at https://github.com/zhongchenzhao/FlashEQLinear.