arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07086cs.LGcs.CV

条件化初始化用于注意力机制

Conditioned Initialization for Attention

Hemanth Saratchandran, Simon Lucey

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出条件化初始化方法,通过改善注意力层谱性质降低雅可比条件数,加速收敛并提升泛化,且易于集成到多种Transformer架构中。

中文摘要 AI 辅助

Transformer是现代机器学习中的主导架构,为视觉、语言及其他领域的应用提供动力。其成功核心在于注意力层,其中查询、键和值矩阵决定了令牌依赖关系如何被捕获。尽管大量工作集中于扩展和优化Transformer,但相对较少关注查询、键和值的权重如何初始化。常见做法依赖于随机初始化或诸如模仿初始化(从收敛模型中模仿权重模式)和权重选择(从教师模型转移权重)等替代方案。在本文中,我们认为初始化可以引入一种优化偏差,从根本上塑造训练动态。我们提出条件化初始化,一种原则性方案,通过初始化注意力权重来改善注意力层的谱性质。理论上,我们证明条件化初始化可能降低注意力雅可比矩阵的条件数,从而带来更稳定的优化。实证上,它加速收敛并提升跨多种应用的泛化能力,凸显了条件化作为提升Transformer性能的关键但尚未充分探索的领域。重要的是,条件化初始化易于应用,并能无缝集成到广泛的Transformer架构中。

英文摘要

Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their success lies the attention layer, where the query, key, and value matrices determine how token dependencies are captured. While considerable work has focused on scaling and optimizing Transformers, comparatively little attention has been paid to how the weights of the queries, keys and values are initialized. Common practice relies on random initialization or alternatives such as mimetic initialization, which imitates weight patterns from converged models, and weight selection, which transfers weights from a teacher model. In this paper, we argue that initialization can introduce an optimization bias that fundamentally shapes training dynamics. We propose conditioned initialization, a principled scheme that initializes attention weights to improve the spectral properties of the attention layer. Theoretically, we show that conditioned initialization can potentially reduce the condition number of the attention Jacobian, leading to more stable optimization. Empirically, it accelerates convergence and improves generalization across diverse applications, highlighting conditioning as a critical yet underexplored area for advancing Transformer performance. Importantly, conditioned initialization is simple to apply and integrates seamlessly into a wide range of Transformer architectures.

发表机构

  • Australian Institute for Machine Learning(澳大利亚机器学习研究所)
  • Adelaide University(阿德莱德大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑