arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

论视觉Transformer中注意力与前馈网络分离的必要性

On the Necessity of Attention-FFN Split in Vision Transformers

Junhyeok Kim, Jinyeong Kim, Jae Wan Park, Seong Jae Hwang

arXiv 2610.10303首次发表:更新:

发表机构

Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AttenFeed模块和统一视觉Transformer(uViT),研究ViTs中注意力-FFN分离的必要性,发现该二分法在小规模模型上因刚性参数分配而损害性能,并提供理论洞见。

AI 中文摘要

标准Transformer架构依赖于一种刚性的模式,即交替使用注意力层和前馈网络(FFN)层。尽管这种架构被广泛采用,但这种严格分离所施加的归纳偏置尚未得到系统性的检验。在本工作中,我们研究了视觉Transformer(ViTs)中注意力-FFN二分法的必要性。为便于分析,我们引入了AttenFeed模块,这是一个统一组件,集成了注意力和FFN的功能特性。基于该模块,我们设计了统一视觉Transformer(uViT),用一系列AttenFeed模块取代了传统的交替注意力-FFN结构。然后,我们将uViT作为对照组,放宽标准ViT的注意力-FFN二分法,并在多个数据集和模型规模上系统比较这两种模型。我们的实验表明,由于ViTs的刚性参数分配,注意力-FFN二分法在较小模型规模下会阻碍性能。AttenFeed模块和uViT可作为理解注意力-FFN结构的新分析工具,并为传统ViTs的启发式设计架构提供理论见解。

英文摘要

The standard Transformer architecture relies on a rigid pattern that alternates Attention and Feed-Forward Network (FFN) layers. Despite its widespread adoption, the inductive bias imposed by this strict separation has not been systematically examined. In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs). To facilitate this analysis, we introduce the AttenFeed module, a unified component that integrates the functional properties of both Attention and FFN. Based on this module, we devise the unified Vision Transformer (uViT), which replaces the conventional alternating Attention-FFN structure with a sequence of AttenFeed modules. We then use uViT as a control group that relaxes the Attention-FFN dichotomy of the standard ViT and systematically compare the two models across multiple datasets and model scales. Our experiments reveal that the Attention-FFN dichotomy can hinder performance at smaller model scales due to the rigid parameter allocation of ViTs. The AttenFeed module and uViT serve as new analytical tools for understanding the Attention-FFN structure and offer theoretical insights into the heuristically designed architecture of conventional ViTs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑