arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

能力涌现可以被预测:按种子、提前、带校准区间、认证误报及盲预注册门控

Capability Emergence Can Be Forecast: Per-Seed, In Advance, With Calibrated Intervals, Certified False Alarms, and a Blind Pre-Registered Gate

Gunner Levi Howe

arXiv 2609.19000首次发表:更新:

AI 中文总结

本研究证明能力涌现可被预测:通过前一个token头形成时间提前预测归纳头涌现,带校准区间和认证误报,并验证了乘法时间规律。

AI 中文摘要

涌现能力被广泛视为不可预测的:损失平滑改善而能力突然出现。先前的工作提供了早期预警指标,但从未将其作为预测进行评分:没有在受控误报率下的提前时间,没有校准,没有负例,没有盲测。我们提供了这一纪律,并表明在grokking模型系统和小型语言模型中,涌现时间可以按每次运行、提前、以校准的不确定性进行预测。在30个相同配置的Transformer中,前一个token头(previous-token head)的形成时间预测每个种子的归纳头(induction head)涌现,Spearman rho=0.977,中位提前975步(约占训练的15%);一个最佳情况损失规则以50步提前(一个即时预测)平了排名。共形区间覆盖了15/15个留出种子,冻结规则通过了两个从未见过的配置上的盲预注册门控(10/10和9/10覆盖率)。随后,一个陷阱语言阶梯攻击了我们自己的规则作为预注册:在前一个token上下文为任务本身付费的情况下,裸前体在10/10个能力阻断运行上误报,而机制组合合取在两个语言类别中均被认证(0误报),并以rho=1.000计时涌现。最后,一项间隙起源研究打破了固定偏移(lr和batch都将间隙移动约2.3倍;没有外部时钟拥有它),并揭示了其下的规律:在80个有效锚点运行中,锚点在时间到涌现的0.843处触发——t_event ≈ 1.19 x t_anchor——这个乘法规则在第三个未见配置上通过了其自身的盲门控(5/5)。误报针对33个人工制造的负例进行了认证。前体在3个公共模型家族(Pythia、OLMo、OLMo-2;7个套件)中领先,其中OLMo-2在1B token时显示前体已形成而能力缺失。四个预注册的终止标准被触发并已报告。每次冻结都在公共提交链中先于其数据。

英文摘要

Emergent capabilities are widely treated as unpredictable: loss improves smoothly while abilities appear abruptly. Prior work offers early-warning indicators but never scores them as forecasts: no lead time at controlled false-alarm rate, no calibration, no negatives, no blind tests. We supply that discipline and show that, in grokking model systems and small language models, emergence timing is forecastable per run, in advance, with calibrated uncertainty. Across 30 transformers at identical configuration, the formation time of the previous-token head forecasts each seed's induction-head emergence at Spearman rho=0.977 with median lead 975 steps (~15% of training); a best-case loss rule ties the ranking with 50-step lead (a nowcast). Conformal intervals covered 15/15 held-out seeds, and the frozen rule passed blind pre-registered gates on TWO never-seen configurations (10/10 and 9/10 coverage). A trap-language rung then attacked our own rule as pre-registered: where previous-token context pays for the task itself, the bare precursor false-alarms on 10/10 capability-blocked runs, while the mechanism-composed conjunction is certified in both language classes (0 false alarms) and times emergence at rho=1.000. Finally, a gap-origin study broke the fixed offset (both lr and batch move the gap ~2.3x; no external clock owns it) and revealed the law beneath: across 80 valid-anchor runs the anchor fires at 0.843 of time-to-emergence -- t_event ~= 1.19 x t_anchor -- and this multiplicative rule passed its own blind gate (5/5) at a third unseen configuration. False alarms are certified against 33 manufactured negatives. The precursor leads across 3 public model families (Pythia, OLMo, OLMo-2; 7 suites), with OLMo-2 at 1B tokens showing precursor formed, capability absent. Four pre-registered kill criteria fired and are reported. Every freeze precedes its data in a public commit chain.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑