AI 中文总结
本研究训练26M参数Transformer模型,通过合成语言构造共现规则使模型依赖线索无法区分,发现回路前的时间等可重复规律,明确机制归因于数据的适用标准。
AI 中文摘要
当上下文对一个事实断言两个值时,模型会依赖近期性、重复性、位置这类线索,但自然数据中这些线索极少出现不一致,因此模型行为无法揭示其依赖的线索。我们在一种合成语言上训练了参数规模为26M的Transformer模型,其中近期性和稀有性完全共现,并通过最小的因果编辑将一条线索反转,同时保持真值、标记数量和答案位置固定。全部75次运行的准确率均达到≥0.999,包括那些简单启发式方法失效的情况,因此没有保留的评估能区分这些模型。在干预下,每个单元的读数并不重复:25个单元中有13个在三个随机种子间的符号分数差异超过0.3,最大差异为0.879,而标准误差为0.025。该构造预测了这一结果——共现规则让目标函数对它们无差异,且方差由每次比较释放的优化量决定。可重复的是时间:从位置捷径的逃逸有一个闭式上限,且随冗余度单调变化。在逃逸前探测时,75次运行中有32次的归因在准确率不变的情况下符号反转,且对回路形成的门控是必要但非充分条件。语料库确定了机制何时出现,而非哪个机制出现——这是一个关于何时可将机制归因于数据的标准,而我们的构造使不可用的情况变得精确。
英文摘要
When a context asserts two values for one fact, a model commits to a cue -- recency, repetition, position -- but natural data rarely makes these disagree, so behavior cannot reveal which. We train 26M-parameter transformers on a synthetic language where recency and rarity are exactly coextensive, and separate them with a minimal causal edit that inverts one cue while holding the truth, token count and answer position fixed. All 75 runs reach accuracy >= 0.999, including where the trivial heuristic fails, so no held-in evaluation distinguishes them. Under intervention the per-cell readout does not replicate: 13 of 25 cells differ by more than 0.3 in sign fraction across three seeds, the largest by 0.879 against a standard error of 0.025. The construction predicts this -- coextensive rules leave the objective indifferent between them -- and the variance is ordered by how much of the optimization each comparison releases. What replicates is timing: escape from a positional shortcut with a closed-form ceiling, monotone in redundancy. Probed before that escape, attribution reverses sign in 32 of 75 runs at unchanged accuracy, and gating on circuit formation is necessary but not sufficient. The corpus fixes when a mechanism appears, not which one -- a criterion for when mechanistic attribution to data is available at all, and our construction makes the unavailable case exact.
Comments37 pages, 3 figures, 17 tables