RoLA:面向高效扩散Transformer的旋转位置低秩线性注意力
RoLA: Rotary-Positioned Low-Rank Linear Attention for Efficient Diffusion Transformers
- Peking University(北京大学)
- Tsinghua University(清华大学)
- University of Electronic Science and Technology of China(电子科技大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对扩散Transformer中自注意力二次方复杂度问题,提出RoLA,通过旋转位置低秩线性注意力保持全局聚合,在90%稀疏度下实现2.63倍推理加速。
中文摘要 AI 辅助
扩散Transformer(DiTs)在视频生成质量上表现出色,但其密集的时空自注意力计算量随序列长度呈二次方增长,并迅速成为推理的主要瓶颈。稀疏低秩混合方法通过结合局部稀疏分支与全局压缩分支来缓解这一成本。在配备3D旋转位置编码(RoPE)的视频DiTs中,全局分支面临结构兼容性问题:当RoPE在非线性特征映射之前应用时,旋转与非线性通常不可交换,这使得在保持相对旋转几何特性的同时维持查询无关的线性摘要变得困难。现有工作通常通过以坐标条件代理或可学习的绝对位置模块替代真正的跨令牌全局聚合来回避此问题。这些折中方案可能有效,但它们从绝对坐标近似相对衰减,并引入了额外的位置参数。我们提出RoLA,一种旋转位置低秩线性注意力分支,它保持真正的跨令牌聚合,同时与可复用的线性摘要兼容。该设计将RoPE应用于非线性低秩特征映射之外,并重用预训练旋转调度中与低秩瓶颈匹配的截断子集。这产生了一个线性时间的低秩全局分支,其相对位置行为由设计保证,且无需额外位置参数;完整的稀疏-低秩模块仍包含固定稀疏性的稀疏分支。在开源视频DiTs上的实验表明,所提方法在90%稀疏度下保持有竞争力的生成质量,同时在Wan2.1-14B(720p,81帧,在NVIDIA H100 GPU上测量)上实现了2.63倍的端到端推理加速。
英文摘要
Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequence length and quickly becomes the dominant inference bottleneck. Sparse low-rank hybrids alleviate this cost by combining a local sparse branch with a global compressed branch. In video DiTs equipped with 3D Rotary Position Embeddings (RoPE), the global branch faces a structural compatibility issue: when RoPE is applied before a nonlinear feature map, the rotation and nonlinearity generally do not commute, making it difficult to keep a query-independent linear summary while preserving relative rotary geometry. Existing work often sidesteps this issue by replacing genuine cross-token global aggregation with coordinate-conditioned surrogates or learnable absolute positional modules. These compromises can be effective, but they approximate relative decay from absolute coordinates and introduce extra positional parameters. We propose \textbf{RoLA}, a rotary-positioned low-rank linear-attention branch that keeps genuine cross-token aggregation while remaining compatible with a reusable linear summary. The design applies RoPE \emph{outside} the nonlinear low-rank feature map and reuses a truncated subset of the pre-trained rotary schedule matched to the low-rank bottleneck. This yields a linear-time low-rank global branch with relative positional behavior by design and no additional positional parameters; the full sparse--low-rank module still includes the fixed-sparsity sparse branch. Experiments on open-source video DiTs show that the resulting method remains competitive in generation quality at 90\% sparsity while achieving 2.63$\times$ end-to-end inference speedup on Wan2.1-14B (720p, 81 frames, measured on an NVIDIA H100 GPU).