arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

复制相同,蒸馏差异:初始化线性视觉Transformer

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Shiqi Huang, Min Kass Chong, Wahyu Wiratama, Peng Hu, Chen Gong, Wu Liu, Xi Peng, Chun Jian Ho, Hongyuan Zhu

arXiv 2609.35745首次发表:更新:

发表机构

Institute of Advanced Intelligence and Computing (IAIC), A*STAR; Nanyang Technological University; ST Engineering Geo-Insights; Sichuan University; Shanghai Jiao Tong University; University of Science and Technology of China(先进智能与计算研究所(IAIC),新加坡科技研究局; 南洋理工大学; 新科工程地理信息公司; 四川大学; 上海交通大学; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究线性ViT的初始化,发现复制MLP权重并蒸馏注意力可有效迁移Softmax ViT预训练权重,使线性ViT性能匹敌甚至超越Softmax版本。

AI 中文摘要

线性视觉Transformer(ViT)旨在用线性复杂度的注意力算子替换Softmax ViT中的注意力,以实现更高效的特征路由,但它们需要从头预训练,且通常性能不如原始的Softmax版本。如何高效且有效地初始化线性ViT仍不清楚。在这项工作中,我们明确提出问题:鉴于大多数基础ViT基于主流的Softmax注意力构建,线性ViT能否从其预训练权重中受益?近期关于注意力迁移的研究表明,注意力是Softmax ViT之间有效的可迁移组件,暗示仅注意力就足以实现这种复用。然而,我们发现Softmax到线性迁移的情况恰恰相反。注意力权重是算子特定的:直接复制它们几乎没有帮助,有时甚至比随机初始化更差。相反,通过适当的损失设计进行蒸馏,可以恢复注意力的特征路由行为,使线性ViT缩小差距甚至匹配Softmax ViT。相比之下,携带学习表示的MLP权重是算子无关的:它们可以通过简单直接复制来迁移,这已经带来了预训练权重的大部分好处。因此,复制MLP可以作为Softmax到线性迁移的有效基础:结合蒸馏的注意力,线性ViT最终缩小剩余差距甚至超越Softmax ViT。这些发现在各种线性ViT变体、不同模型规模和多样数据集上保持一致。我们希望这项研究能加深对跨注意力算子复用预训练权重的理解:复制相同之处,蒸馏不同之处,以恢复跨Softmax到线性边界的好处。

英文摘要

Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑