发表机构
Salk Institute(索尔克研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过消融实验发现,提高国际象棋Transformer的技能输入水平会单调地将计算推向更深层,尤其对马叉等战术影响显著,揭示了条件输入如何重新分配计算资源。
AI 中文摘要
国际象棋涉及在确定性环境中的复杂推理,这使其成为研究Transformer内部计算机制的有用场景。Maia-3国际象棋Transformer将Elo(一种竞技国际象棋技能的衡量标准)作为预训练网络的输入,因此我们可以在不改变网络权重的情况下改变网络所依赖的技能水平。在此,我们研究转动这一技能旋钮如何影响自注意力。通过消融从700到2500每个Elo等级下的每个注意力头,我们发现:1)提高技能水平会单调地将计算因果质心推向更深层,且这一现象在我们测量的每一种棋子和走法类型中都成立;2)对于特定战术,尤其是马叉,深度迁移远大于其他走法类型;3)这种迁移表现为更深的头被招募用于更专门化的计算,而一个共享的浅层头保持大致恒定的贡献。这些结果可能有助于阐明条件输入如何在更大的Transformer中重新分配计算。
英文摘要
Chess involves complex reasoning in a deterministic environment, which makes it a useful setting for studying the mechanisms of computation inside transformers. The Maia-3 chess transformer takes Elo, a measure of competitive chess skill, as an input to the pre-trained network, so we can vary the skill the network is conditioned on with no change to its weights. Here we investigate how turning this skill dial affects self-attention. Ablating every attention head at every Elo from 700 to 2500, we find 1) increasing skill pushes the causal center of mass of the computation deeper, monotonically, for every chess piece and move type we measured; 2) the depth migration is much greater for specific tactics, especially knight forks, than for other move types; 3) the migration consists of deeper heads getting recruited for more specialized computations while one shared shallow head keeps a roughly constant contribution. These results may shed light on how conditioning inputs redistribute computation in larger transformers.
Comments19 pages, 12 figures. Code and data: https://github.com/David-31415/maia-depth-migration. Built with chessformer-lens library: https://github.com/chessformer-lens/chessformer_lens