数据高效语言建模:从前沿推进到原则指导的模型改进
Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement
浏览论文内容
中文总结 AI 辅助
本研究通过三阶段自主研究,在数据受限下提出可测试的数据高效学习原则,并据此改进模型,使总体得分从42.02提升至42.25,实现研究过程的递归自我改进。
中文摘要 AI 辅助
从有限的文本中学习要求模型利用上下文、泛化到新输入并保留有用的能力。Qiushi Engine在BabyLM 2026 Strict-Small上开展了一项长期、端到端的自主研究项目,语料规模为1000万词,累计词展示量为1亿次。三个阶段连接了前沿推进、原则发现和原则指导的模型改进。第一阶段结合了紧凑重述、预算再投资和残差增量学习来构建前沿模型。第二阶段发现,精确重复和对齐重述会产生不同的上下文使用模式,这取决于目标关系和预测窗口。在受控任务中,恢复熟悉的表现并不能确保未见输入仍能使用已学习的计算。这些发现支持一个可测试的数据高效学习原则:围绕预测所需的上下文依赖组织经验;分别设计可见信息、监督和保留;测试学习、泛化和保留。第三阶段保留了源文本,掩盖了更多局部线索,监督了选定的目标,并保留了对通常被掩盖输入的预测。来自同一父模型的两个续训种子在完整的九指标聚合上优于普通续训。总体得分在两代中从42.02上升到42.25;第二代在2026年9月8日的公开Strict-Small快照中取得了最高总体得分。进一步的研究涉及压缩、关系锚点、共享表示和测量。模型可在Hugging Face上获取;代码和研究记录随GitHub仓库提供。总之,这些阶段展示了研究RSI:研究过程的递归自我改进。科学理解和方法创新改变了后续的问题和设计;新实验对其进行测试和完善。
英文摘要
Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, within 10 million corpus words and 100 million cumulative word presentations. Three stages connected frontier advancement, principle discovery, and principle-guided model improvement. Stage I combined compact restatements, budget reinvestment, and residual incremental learning to build a frontier model. Stage II found that exact repetition and aligned restatement produce different patterns of context use, depending on target relations and prediction windows. In controlled tasks, recovering familiar performance did not ensure that unseen inputs could still use learned computations. These findings support a testable data-efficient learning principle: organize experience around the contextual dependencies needed for prediction; separately design visible information, supervision, and preservation; test learning, generalization, and retention. Stage III retained source text, masked more local clues, supervised selected targets, and preserved predictions on ordinarily masked inputs. Two continuation seeds from the same parent outperformed ordinary continuation on the complete nine-metric aggregate. Overall rose from 42.02 to 42.25 across two generations; the second achieved the highest Overall in the public Strict-Small snapshot of 8 September 2026. Further studies addressed compression, relational anchors, shared representations, and measurement. Models are available on Hugging Face; code and research records accompany the GitHub repository. Together, these stages illustrate Research RSI: recursive self-improvement of the research process. Scientific understanding and method innovations change subsequent questions and designs; new experiments test and refine them.
发表机构
- College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息与电子工程学院)
机构由 AI 辅助整理,请以论文原文为准。