arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小到足以知晓一切:完全可枚举Transformer作为延迟泛化科学的仪器

Small Enough to Know Everything: The Fully-Enumerable Transformer as an Instrument for the Science of Delayed Generalization

Yoshiyuki Ootani

arXiv 2609.20166首次发表:更新:

AI 中文总结

本研究将完全可枚举任务上的小型Transformer作为科学仪器,通过预注册的守恒研究验证三条任务侧定律在12K至50M参数规模下的普适性,发现两条守恒、一条系统性变形,证明该范式可揭示延迟泛化的可测量规律。

AI 中文摘要

在完全可枚举任务上训练的小型Transformer在grokking研究中占据了一个不寻常的位置:每个输入都可以被评估,每个泛化上限都可以被精确计算,并且数百个种子只需几分钟。我们认为这一机制是一种科学仪器,具有近似设置无法提供的四种能力:(a)精确的、可证伪的泛化上限;(b)任务手术,可在可证明固定所有其他变量的同时操纵一个结构变量;(c)直接观察每个权重;以及(d)跨多个种子的生存时间统计,将“不grok”重新表述为截尾观测。明显的反对意见是,在10^4参数下表征的定律可能在其之外毫无意义。我们通过一项预注册的守恒研究来回答这一问题:在12K参数下确立的三条任务侧定律——可恢复性上限定律、角色冲突延迟定律和权重衰减响应定律——在12K、1M和50M参数(跨度4000倍;360次运行外加44次运行的对照组)下,以相同的从头开始协议重新测量。上限定律和延迟定律得以守恒(0/144个Holm校正上限违规;每个规模下Spearman rho >= 0.75,置换p < 1e-4),而权重衰减定律则系统性变形,随规模增大而变得更陡峭。预注册的对照表明,50M角色冲突缺陷在学习率调整和预算增加三倍后依然存在。守恒性测试针对数据收集前冻结的标准进行,而一条定律的变形表明该测试本可能失败。这些结果使完全可枚举Transformer成为延迟泛化任务侧定律的模式生物:它所精确测量的,较大模型大体上遵循——而在它们偏离之处,偏离本身也是有规律且可测量的。

英文摘要

Tiny transformers trained on fully-enumerable tasks occupy an unusual position in the study of grokking: every input can be evaluated, every generalization ceiling can be computed exactly, and hundreds of seeds cost minutes. We argue this regime is a scientific instrument with four capabilities that approximate settings cannot offer: (a) exact, falsifiable generalization ceilings; (b) task surgery that manipulates one structural variable while provably fixing all others; (c) direct observation of every weight; and (d) survival-time statistics over many seeds that recast "does not grok" as a censored observation. The obvious objection is that laws characterized at 10^4 parameters may not mean anything beyond them. We answer it with a preregistered conservation study: three task-side laws established at 12K parameters -- a recoverability-ceiling law, a role-conflict delay law, and a weight-decay response law -- are re-measured under an identical from-scratch protocol at 12K, 1M, and 50M parameters (a 4,000x span; 360 runs plus a 44-run control arm). The ceiling law and the delay law are conserved (0/144 Holm-corrected ceiling violations; Spearman rho >= 0.75 at every scale, permutation p < 1e-4), while the weight-decay law deforms systematically, steepening with scale. Preregistered controls show the 50M role-conflict deficit survives learning-rate adjustment and a tripled budget. Conservation was tested against criteria frozen before data collection, and the third relationship fails that test even at the original scale, which is what tells us the test could have failed. The series scales the model while the tasks stay enumerable, so what it licenses is specific: the quantities this regime computes from a task keep predicting once the model has outgrown the regime itself.

Comments10 pages, 2 figures. v2 narrows the central claim: the study scales the model while every task stays enumerable, so it shows that enumerated task-side quantities keep predicting once the model outgrows the instrument's regime, not that they survive the loss of enumerability. The weight-decay relationship is demoted; data and figures are unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑