发表机构
Australian Institute for Machine Learning; Adelaide University(澳大利亚机器学习研究所; 阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对 Transformer 在布尔外推任务上的系统性错误泛化问题,提出直接初始化注意力掩码的方法,该方法可保留结构先验,在布尔推理等任务上实现优异外推效果。
AI 中文摘要
Transformer 不仅在一些布尔外推任务上缺乏数据,还会以系统性错误的方式进行泛化。近期关于未见数据泛化的研究表明,尽管 Transformer 能拟合观测到的领域,但它们通常会依据更简单的最小次数插值器而非真实目标函数进行外推。这些布尔任务并非实际应用,而是用于理解 Transformer 归纳偏置的受控压力测试。本研究探究是否可通过向注意力机制注入显式结构先验来修正该失败模式。现有结构化初始化方法通过选择查询(query)和键(key)投影使相似度分数近似期望注意力模式,间接改变 Transformer 的归纳偏置。然而,我们发现将这些基于查询-键(QK)的先验应用于布尔外推时,会在训练过程中被快速覆盖,无法改变学习到的外推规则。我们提出一种更简单的替代方案:直接初始化加性注意力掩码。与用于因果或局部注意力的标准硬掩码不同,我们的掩码是从任务级交互结构初始化的有限、可学习的注意力对数偏差。这将结构先验与依赖内容的注意力分数分离,使其在整个优化过程中得以保留。在布尔推理任务上,基于掩码的初始化实现了近乎完美的外推,而普通 Transformer 和基于 QK 初始化的 Transformer 仍受默认归纳偏置的限制。相同机制还可提升少样本算术性能,并在视觉和语言基准上保持竞争力。这些结果表明,注意力掩码不仅可作为架构约束,还可作为在 Transformer 中编码持久归纳偏置的简单载体。
英文摘要
Transformers do not merely lack data on some Boolean extrapolation tasks; they generalize in a systematically wrong way. Recent work on generalization on the unseen has shown that, despite fitting the observed domain, Transformers often extrapolate according to a simpler minimum-degree interpolator rather than the true target function. These Boolean tasks are not practical applications, but controlled stress tests for understanding Transformer inductive bias. We ask whether this failure mode can be corrected by injecting explicit structural priors into attention. Existing structured-initialization methods alter Transformer inductive bias indirectly, by choosing query and key projections whose similarity scores approximate a desired attention pattern. However, we find that when applied to Boolean extrapolation, these QK-based priors can be rapidly overwritten during training and fail to change the learned extrapolation rule. We propose a simpler alternative: initialize the additive attention mask directly. Unlike standard hard masks used for causality or locality attention, our mask is a finite, learnable attention-logit bias initialized from task-level interaction structure. This separates the structural prior from content-dependent attention scores, allowing it to persist throughout optimization. On Boolean reasoning tasks, mask-based initialization achieves near-perfect extrapolation where vanilla and QK-initialized Transformers remain trapped by the default inductive bias. The same mechanism also improves low-data arithmetic performance and remains competitive on vision and language benchmarks. These results show that attention masks can serve not only as architectural constraints, but as a simple substrate for encoding persistent inductive bias in Transformers.