注意力函数作为内在归纳偏置:模型在新颖情境中的行为差异
Attention Function as an Intrinsic Inductive Bias: How Models' Behavior Diverges in Novel Contexts
- Georgia Institute of Technology(佐治亚理工学院)
- Seoul National University Hospital(首尔大学医院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出MoFA,通过固定注意力头中softmax与sigmoid比例作为内在归纳偏置,发现该比例在分布内影响甚微,但在零样本分布偏移下显著影响模型行为,并沿文本类型轴解释78.3%的方差。
AI中文摘要:
发展心理学认为,某些先验是在婴儿获得经验之前就给予的,而非从数据中归纳而来,并且这些先验的影响在强约束条件下被抑制,但在弱约束条件下会重新显现。我们探讨Transformer是否遵循类似原理:给予注意力头的激活函数能否作为内在归纳偏置?我们提出了函数混合注意力(MoFA),一种对多头注意力的无参数修改,在训练前固定softmax和sigmoid头的比例。在五种比例、一个124M参数的GPT-2模型和五个随机种子下,我们发现该给定比例在分布内影响甚微——对于中等混合,比例间的差异在统计上可忽略,即使在极端情况下差异也很小——但在零样本分布偏移下,其影响在15个分布外领域重新显著出现。在几个领域上,比例间的困惑度差距扩大了一个数量级以上,且最佳比例沿着单一领域结构轴,将短小非正式文本(偏好softmax)与技术性长文(偏好sigmoid)分开,这解释了领域响应中78.3%的方差。这种重组在头部层面可见:sigmoid头的注意力熵随着其比例增加而加速下降,而softmax头的响应较为温和,导致两种头型之间产生一致的分工。我们的结果表明,激活函数的选择起到给定先验的作用,其影响在分布内被掩盖,而在分布外重新显现。
英文摘要:
Developmental psychology holds that certain priors are given to infants prior to experience rather than induced from data, and that the influence of such priors is suppressed under strong, well-constrained conditions but reasserts itself under weak ones. We ask whether an analogous principle holds for the Transformer: can the activation function given to attention heads serve as an intrinsic inductive bias? We propose Mixture of Function Attention (MoFA), a parameter-free modification to multi-head attention that fixes a ratio of softmax and sigmoid heads before training. Across five ratios, a 124M-parameter GPT-2 model, and five seeds, we find that this given ratio has little effect in-distribution -- differences between ratios are statistically negligible for moderate mixtures and remain small even at the extremes -- but its influence re-emerges sharply under zero-shot distribution shift across 15 out-of-distribution domains. Perplexity gaps between ratios widen by more than an order of magnitude on several domains, and the best-performing ratio tracks a single axis of domain structure, separating short, informal text (softmax-favoring) from technical, long-form text (sigmoid-favoring), that explains 78.3% of the variance in domain response. This reorganization is visible at the head level: sigmoid heads show an accelerating drop in attention entropy as their ratio increases, while softmax heads respond more modestly, yielding a consistent division of labor between the two head types. Our results suggest that activation choice functions as a given prior whose influence is masked in-distribution and re-emerges out-of-distribution.