发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过转换博弈定位 Transformer 中记忆到泛化的转变,发现分布式效用增益由第 1 块 MLP 中介,grokking 本质是分布式电路的谱重编码而非模块切换。
AI 中文摘要
在 Transformer 中,从记忆到泛化的转变在功能上体现在何处?我们引入了转换博弈——行为对齐的精确激活博弈,并配以配对的非泛化对照——并发现了带有前瞻性第 0 块注意力偏置的分布式效用增益;选定的二阶模态在替换博弈中占据了其加法对比度的 67% 至 92%,而一项不相交的精确路径研究证实,在 12/12 对中,第 1 块 MLP 比所有其他测试的下游路径更能中介其效果。更尖锐的“MLP 记忆,注意力泛化”预测反而逆转(在记忆锚点处为 -0.331 比特/示例;预测方向上为 0/12),而路由起始、全局秩坍缩以及一个素数不变架构脊也均失败,这表明此处的 grokking 是现有分布式电路的谱重编码,而非模块切换。
英文摘要
Where in a Transformer is the change from memorization to generalization functionally expressed? We introduce Transition Games--behavior-aligned exact activation games with paired non-generalizing controls--and find distributed utility gain with a prospective block-0 attention bias; selected degree-two modes account for 67--92% of its addition contrast across replacement games, and a disjoint exact path study confirms that block-1 MLP mediates more of their effect than all other tested downstream paths in 12/12 pairs. The sharper "MLP memorizes, attention generalizes" prediction instead reverses (-.331 bits/example at the memory anchor; 0/12 in the predicted direction), while routing onset, global rank collapse, and a prime-invariant architecture ridge also fail, identifying grokking here as spectral recoding of an existing distributed circuit rather than a module switch.
Comments27 pages, 10 figures. Code and data: https://github.com/cheshireyang/where-grokking-happens/releases/tag/v0.3.2