预训练期间的知识获取?辅助视角让大语言模型学习得更好
Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views
浏览论文内容
中文总结 AI 辅助
本研究通过受控实验证实,辅助视角(知识的重构形式)是大语言模型预训练成功的关键因素,可提升学习效果,解释了数据多样性的重要性。
中文摘要 AI 辅助
目前对大语言模型(LLMs)在预训练期间如何获取知识的理解仍存在缺口。我们假设知识的重构形式——辅助视角,对学习具有因果层面的帮助。我们设计受控实验来分离验证该假设:首先,确认重复是知识获取的必要条件,并明确仅在较小批次规模下,改写才会有帮助;其次,在固定 token 预算的情况下,将原本用于文档重复的 token 分配给辅助视角,会提升学习效果,甚至在事实回忆任务中也成立;第三,辅助视角的有效性不依赖于生成它们的教师模型的强弱;第四,我们识别出在存在先验知识缺口时,有助于学习的知识类型:上下文知识与基础知识;最后,我们通过分层偏差与压缩机制,探究这些效应的具体表现机制。综上,我们的研究结果表明,大型预训练语料库中自然产生的知识的辅助表征,是预训练成功的关键因素,也为数据多样性为何重要提供了合理的解释。
英文摘要
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Finally, we examine how these effects manifest mechanistically via layer-wise biases and compression. Together, our findings suggest that auxiliary representations of knowledge, which arise naturally in large pre-training corpora, are a key factor in the success of pre-training and offer a plausible explanation for why data diversity matters.
发表机构
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。