现代Transformer是隐式混合体:从功能分化到原则性混合架构设计
Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design
查看机构详情
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究基于Transformer头级功能组织提出头级混合架构HwH,结合NoPE的FA与LA,在FA-LA比低于1:3时,提升了检索与零样本长上下文外推能力,保留了原有性能。
中文摘要 AI 辅助
结合全注意力(Full Attention, FA)与线性注意力(Linear Attention, LA)的混合架构日益受到关注,但其分配方式仍依赖启发式方法。我们在基于旋转位置嵌入(RoPE)的Transformer所学习的头级功能组织中寻求基于证据的依据。行为探测未得出完整分类,因此我们提出两个干预指标:RoPE频率重要性分数(RoPE Frequency Importance Score, RFIS),用于衡量每个频率对注意力头注意力分布的影响;RoPE位置依赖性(RoPE Positional Dependence, RPD),用于分离对旋转位置调制的依赖性。在Qwen3系列模型和Llama3.1上,RFIS表明且RPD验证了由显著中低频带分隔的检索头与位置头的完整分类。受控Transformer显示,该边界遵循训练长度位置尺度;我们将其命名为全局位置带(Global Positional Band, GPBand)。该分析指出了零样本长度外推失败的潜在原因,并得出两个原则:位置建模应仅在局部进行,通过与位置无关的检索实现全局访问;且两种功能应在头级别分配,采用层特定分配方式。我们将其实例化为头级混合架构(Head-wise Hybrid Architecture, HwH),使用无旋转位置嵌入(NoPE)的FA进行全局检索,LA用于局部位置建模。在FA与LA比例低于1:3的情况下,HwH在保留强大语言建模和常识推理能力的同时,提升了检索性能,并较Transformer、LA及层级混合基线显著增强了零样本长上下文外推能力。消融实验验证了两项原则及组件作用,凸显原则性混合架构设计是未来基础模型的有前景路径。
英文摘要
Hybrid architectures combining Full Attention (FA) and Linear Attention (LA) are increasingly prominent, yet their allocation remains heuristic. We seek an evidence-grounded basis in head-level functional organization learned by RoPE-based Transformers. Behavioral probes do not yield a complete taxonomy, so we propose two intervention metrics: RoPE Frequency Importance Score (RFIS), measuring how each frequency affects a head's attention distribution, and RoPE Positional Dependence (RPD), isolating dependence on rotary positional modulation. On Qwen3-series models and Llama3.1, RFIS suggests and RPD verifies a complete taxonomy of retrieval and positional heads separated by a salient mid-low-frequency band. Controlled Transformers show that this boundary follows the training-length positional scale; we term it the Global Positional Band (GPBand). The analysis suggests a potential cause of zero-shot length-extrapolation failure and yields two principles: positional modeling should operate only locally, with global access through position-independent retrieval; and both functions should be assigned at head granularity with layer-specific allocation. We instantiate them in Head-wise Hybrid Architecture (HwH), using NoPE FA for global retrieval and LA for local positional modeling. With an FA-to-LA ratio below 1:3, HwH retains strong language modeling and commonsense reasoning while improving retrieval and substantially strengthening zero-shot long-context extrapolation over Transformer, LA, and a layer-wise hybrid baseline. Ablations validate both principles and component roles, highlighting principled hybrid architecture design as a promising route toward future foundation models.