arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

查询扩展与键特化在Transformer注意力几何中的表现

Query Expansion and Key Specialization in Transformer Attention Geometry

Vidit Gupta, Siddhesh Nadkarni, Mihik Chaudhari, Vinaya Sawant, Prachi Tawde

arXiv 2609.34273首次发表:更新:

发表机构

Dwarkadas J. Sanghvi College of Engineering(德瓦卡达斯·J·桑格维工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过训练不同深度和初始化的小型GPT模型,发现查询有效维度扩展而键收缩,并证明键谱收缩可因果性地锐化注意力,该趋势在早期训练中显著但后期衰减。

AI 中文摘要

查询和键的投影是Transformer架构中注意力机制的核心。尽管它们在数学上是对称的,但在注意力机制中扮演着不同的角色。它们的功能差异是否会影响训练过程中几何发展的问题仍未得到解答。我们通过在字符级WikiText-103上训练小型类GPT Transformer来研究该问题,涉及三种不同的深度(4、6和8层)、三种查询和键的初始化类型以及四个随机种子,共产生36次运行和54条跨种子的平均层轨迹。我们使用参与比追踪这些层的有效维度,发现查询的有效维度扩展而键的有效维度收缩,并且所有研究轨迹中$PR_Q - PR_K$均为正。与注意力相关的是,键的收缩导致$QK^\ op$的谱更窄,注意力权重更尖锐。为确定这种关联是因果关系还是巧合,我们在训练期间直接控制键的谱(跨五个种子):限制其收缩会使注意力以高方向置信度变得尖锐,而将其保持在初始分散水平则使注意力更柔和。额外的词级检查点分析表明,单调的配对对比趋势并非预训练模型家族的普遍现象,而是作为早期训练阶段存在,随后在完整预训练过程中衰减,并且交互秩几何与注意力熵之间的联系在多个模型中仍然可见。

英文摘要

The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric development in training remains unanswered. We investigate the problem through the training of small GPT-like Transformers on character-level WikiText-103 for three different depths (4, 6, and 8 layers), three types of initialization for queries and keys, and four random seeds, resulting in 36 runs and 54 trajectories of average layers across seeds. We track the effective dimensionality of those layers using participation ratios and discover that effective dimension of queries expand while keys shrink, and that $PR_Q - PR_K$ is positive in all trajectories studied. In connection to attention, the shrinking of keys leads to a narrower spectrum of $QK^\top$ and more peaked attention weights. In order to determine if this connection is causal or coincidental, we directly control the spectrum of keys during training across five seeds: restricting it to make it shrink sharpens the attention with high directional confidence, while keeping it constant to the level of initial dispersion makes attention softer. Additional token-level checkpoint analyses show that the monotonic paired-contrast trend is not universal across pretrained families, but survives as an early-training regime that later decays over a full pretraining run, and the link between interaction-rank geometry and attention entropy remains visible in several models.

CommentsAccepted at Asian Conference on Machine Learning 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑