arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过带随机LayerNorm的Transformer处理追踪可区分性

Tracing distinguishability through transformer processing with stochastic LayerNorm

Kieran Murphy

arXiv 2608.30720首次发表:更新:

发表机构

New Jersey Institute of Technology(新泽西理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出带随机LayerNorm的Transformer方法,将表征相似性转化为统计可区分性,通过Bhattacharyya系数追踪区分性,在ViT-S和GPT-2 small实验中验证了其功能基础视角,补充了Transformer可解释性方法。

AI 中文摘要

表征相似性是深度网络分析的基础,但点值表征之间的距离与下游功能并无本质关联:邻近状态可能产生不同行为,而遥远状态可能表现相似。我们转而赋予表征体积,将相似性转化为统计可区分性。重叠的随机表征必然诱导重叠的下游分布,将潜在比较锚定在模型功能中,并借助信息论工具(如数据处理不等式)实现。我们通过对LayerNorm进行轻量修改,在预训练Transformer中实现了这一思路:在每个残差流读取时,对状态进行归一化,添加各向同性高斯噪声,再重新归一化。在蒸馏微调过程中,每个残差流读取对应一个学习到的分配参数,该参数在整个处理栈中分配固定的全局速率预算。所得模型可视为Transformer块在共享全局速率预算下以学习到的有限精度读取残差流。利用Bhattacharyya系数,我们追踪了哪些反事实区分通过MLP块保留,或选择性暴露于单个注意力头的查询、键和值计算中。在ViT-S和GPT-2 small上的实验表明,连续视觉扰动的深度传播以及与已知注意力基序对齐的token区分的头特异性敏感性。这些结果确立了可区分性作为Transformer计算的功能基础视角,可补充现有的可解释性方法。

英文摘要

Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied to downstream function: nearby states may produce different behaviors, while distant states may behave similarly. We instead give representations volume, turning similarity into statistical distinguishability. Overlapping stochastic representations necessarily induce overlapping downstream distributions, grounding latent comparison in model function and bringing it under information-theoretic tools such as the data-processing inequality. We realize this idea in pretrained transformers through a light-touch modification to LayerNorm: at each residual-stream read, we normalize the state, add isotropic Gaussian noise, and renormalize. During distillation fine-tuning, one learned allocation parameter per residual-stream read distributes a fixed global rate budget across the processing stack. The resulting model can be viewed as transformer blocks reading the residual stream with learned finite precision under a shared global rate budget. Using the Bhattacharyya coefficient, we trace which counterfactual distinctions are preserved through MLP blocks or selectively exposed to the query, key, and value computations of individual attention heads. Experiments on ViT-S and GPT-2 small reveal the depthwise propagation of continuous visual perturbations and head-specific sensitivity to token distinctions aligned with known attention motifs. These results establish distinguishability as a functionally grounded lens on transformer computation that complements existing interpretability approaches.

CommentsCode: https://github.com/murphyka/stoch_layernorm

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑