arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04664cs.CV

FLASHSWIN:利用内存高效注意力解锁Swin视觉Transformer中的大窗口与密集令牌

FLASHSWIN: Unlocking Large Windows and Dense Tokens in Swin Vision Transformers with Memory Efficient Attention

Tushar Kataria, Gerald Sabin, Ponnuswamy Sadayappan, Shireen Y. Elhabian

首次发表
浏览论文内容

中文总结 AI 辅助

FLASHSWIN通过FlashAttention降低窗口注意力内存至O(M²),并用2D RoPE恢复位置信息,实现大窗口密集令牌,提升多个视觉任务性能。

中文摘要 AI 辅助

高分辨率视觉骨干网络长期以来被迫在局部令牌密度与更大的感受野之间做出取舍。分层Swin transformer之所以强加这种折衷,是因为标准窗口注意力在每个窗口内具体化一个$M^2\ imes M^2$的得分矩阵,随着窗口或令牌网格的增长,内存开销达到$O(M^4)$。此外,Swin在注意力得分上逐元素添加学习到的相对位置偏置,要求完整具体化得分矩阵及其梯度。这使得Swin和SwinV2被困在小窗口($M=8,16$)、粗令牌(patch大小为$4\ imes4$,即$p=4$)的机制中,限制了细粒度任务的性能。我们提出FLASHSWIN,用FlashAttention实现替换标准窗口注意力,在不具体化得分矩阵的情况下计算精确的softmax注意力,将每窗口内存从$O(M^4)$降至$O(M^2)$。这实现了更高的令牌密度和更大的感受野,而无需增加内存开销。训练内存在不同窗口大小下保持平坦:在$32\ imes32$窗口下,FLASHSWIN-T仅需$12.4$GB,与$8\ imes8$窗口相同,而SwinV2/V1-T则需要$70/90$GB。然而,直接将FlashAttention应用于Swin会产生权衡:绕过得分矩阵使得无法使用Swin的加性相对位置偏置,从而以空间信息换取内存效率。FLASHSWIN通过窗口局部可学习的2D RoPE恢复位置信息,使大窗口和密集令牌网格既可行又准确。在相同规模下,FLASHSWIN-T优于Swin变体。采用密集令牌和宽窗口($p=2,M=32$)时,同一Tiny模型在ImageNet-1K上达到$84.1\%$,COCO框AP为$44.1$,ADE20K mIoU为$47.28$——分别比$M=16$的SwinV2-T高出$+1.3$、$+5.1$和$+1.82$。在固定$M=32$时,将patch大小减半,边界质量的提升幅度约为mIoU的$3$倍。

英文摘要

High-resolution vision backbones have long been forced to trade away local token density to afford larger receptive fields. Hierarchical Swin transformers impose this compromise because standard windowed attention materializes an $M^2\times M^2$ score matrix per window, incurring $O(M^4)$ memory as windows or token grids grow. Furthermore, Swin adds a learned relative-position bias elementwise to attention scores, requiring full materialization of the score matrix and its gradient. This keeps Swin and SwinV2 trapped in a small-window($M=8,16$), coarse-token regime with patch size $4\times4$ ($p=4$), limiting performance for fine-grained tasks. We introduce FLASHSWIN, which replaces standard windowed attention with a FlashAttention implementation that computes exact softmax attention without materializing the score matrix, reducing per-window memory from $O(M^4)$ to $O(M^2)$. This enables higher token density and larger receptive fields without inflating memory overhead. Training memory is flat across window sizes: at a $32\times32$ window, FLASHSWIN-T requires only $12.4$\,GB, unchanged from $8\times8$, compared to $70/90$\,GB for SwinV2/V1-T. However, applying FlashAttention directly to Swin creates a trade-off: bypassing the score matrix precludes Swin's additive relative-position bias, forfeiting spatial information in exchange for memory efficiency. FLASHSWIN restores position information as window-local learnable 2D RoPE, making large windows and dense token grids both affordable and accurate. At matched scale, FLASHSWIN-T outperforms Swin variants. With dense tokens and wide windows ($p=2,M=32$), the same Tiny model reaches $84.1\%$ ImageNet-1K, $44.1$ COCO box AP, and $47.28$ ADE20K mIoU---gains of $+1.3$, $+5.1$, and $+1.82$ over SwinV2-T at $M=16$, respectively. At fixed $M=32$, halving the patch size yields roughly $3\times$ larger gains in boundary quality than in mIoU.

发表机构

  • Scientific Computing and Imaging Institute, University of Utah(犹他大学科学计算与影像研究所)
  • Kahlert School of Computing, University of Utah(犹他大学卡勒特计算学院)
  • RNET Technologies(RNET技术公司)

机构由 AI 辅助整理,请以论文原文为准。

↑