Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
在LLM预训练中“Grokking”?无需测试即可监控记忆到泛化的过程
机构 * Department of Computer Science, University of Maryland, College Park(计算机科学系,马里兰大学,学院公园)
AI总结 本文研究了LLM预训练中“Grokking”现象,通过分析训练数据路径的动态变化,提出两种新度量标准以监控模型泛化能力,无需额外成本。
Comments Accepted at ICLR 2026