arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16537cs.LGcs.AI

层重要性揭示了Transformer和状态空间模型的什么?

What Does Layer-Importance Reveal About Transformers and State-Space Models?

  • Chandar Research Lab(Chandar 研究实验室)
  • Mila - Quebec AI Institute(Mila - 魁北克人工智能研究所)
  • Université de Montréal(蒙特利尔大学)
  • Polytechnique Montréal(蒙特利尔理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Istabrak Abbes, Nizar Islah, Irina Rish, Sarath Chandar

AI总结:

本研究通过分解层重要性为必要性和可塑性,发现Transformer中两者反对齐而SSM中重叠,且对齐符号预测下游适应行为,影响灾难性遗忘。

AI中文摘要:

Transformer和状态空间模型(SSM)是序列模型的两大主导家族,一个核心的开放问题是:为Transformer建立的分析知识在多大程度上能迁移到SSM。我们通过层重要性这一视角来探讨此问题,层重要性支撑着这两类家族的压缩、选择性微调和可解释性。我们将层重要性分解为两个不同的概念。“必要性”捕捉预训练模型对某一层现有贡献的依赖程度,通过绕过该层导致的损失增加来衡量。“可塑性”捕捉模型在微调过程中吸收新信息的位置,通过任务特定权重更新的幅度来衡量。我们的分析揭示,这两个家族的行为存在根本性差异:在每个评估的、参数高达140亿的残差Transformer中,必要性和可塑性在深度上呈反对齐,而在评估的Mamba风格SSM中,它们指向重叠区域。这种对齐的符号也能预测下游适应行为。在评估的Transformer中,将更新集中在最具可塑性的层会增加灾难性遗忘,而这种层级依赖效应在评估的Mamba风格SSM中消失。

英文摘要:

Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical knowledge built for transformers transfers to SSMs. We address this through the lens of layer importance which underpins compression, selective fine-tuning, and interpretability across both families. We decompose layer importance into two distinct notions. \emph{Necessity} captures how much the pretrained model depends on a layer's existing contribution, measured by the loss increase from bypassing it. \emph{Plasticity} captures where the model absorbs new information during fine-tuning, measured by the magnitude of task-specific weight updates. Our analysis reveals that the two families behave fundamentally differently: in every evaluated residual transformer up to $14$B parameters, Necessity and Plasticity anti-align across depth, whereas in the evaluated Mamba-style SSMs they point to overlapping regions. The sign of this alignment also predicts downstream adaptation behavior. In the evaluated transformers, concentrating updates in the most plastic layers increases catastrophic forgetting, while this tier-dependent effect disappears in the evaluated Mamba-style SSMs.

↑