发表机构
Ellis Institute(埃利斯研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出专为美国小学定制的LITTLECURRICULUM语料库,训练得到受知识边界约束的LittleLearner模型,构建沙盒用于研究模型知识习得,实验证实其注入新知识后不会提升超范围能力。
AI 中文摘要
现代语言模型在异构的网络级文本语料库上进行训练,因此研究其知识与技能习得过程十分困难,因为很难明确模型对相关内容的先前暴露情况。为应对这一挑战,我们推出LITTLECURRICULUM,这是一个专为美国小学材料定制的精选8800亿token预训练语料库,明确排除了五年级以上教授的概念、事实和词汇。在LITTLECURRICULUM上从头训练一个50亿参数的大语言模型(LLM),得到LittleLearner,该模型具备足够的语言能力可用于开放式评估,且其知识与能力边界清晰,可映射到可解释的课程指南。我们发布LITTLECURRICULUM和LittleLearner作为一个发展受限的沙盒,用于研究模型在明确定义的训练范围内如何习得、表征和使用数据。我们通过一组首批实验展示该沙盒的实用性,这些实验通过后训练和上下文学习注入新知识,这些方法让LittleLearner能更好地利用现有知识,但不会提升超出范围的能力。我们的研究结果强调了这种受控环境对未来研究的价值。
英文摘要
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LittleCurriculum, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LittleCurriculum yields LittleLearner, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LittleCurriculum and LittleLearner as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LittleLearner better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.