arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

抽象预训练在困惑度之外的持久影响

Lasting Effects of Abstract Pretraining Beyond Perplexity

Zachary Shinnick, Hemanth Saratchandran, Damien Teney, Anton van den Hengel

arXiv 2609.38764首次发表:更新:

发表机构

Australian Institute for Machine Learning (AIML), Adelaide University; Idiap Research Institute; Metacognition AI(阿德莱德大学澳大利亚机器学习研究所; 伊迪亚普研究所; 元认知AI公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究抽象数据预热对小型语言模型的影响,发现仅用1%预训练代币进行栈操作任务预热,可显著提升多跳问答能力(如MUSIQUE上F1提升3.9点),且效果持久,但收益具有任务特异性和时机依赖性。

AI 中文摘要

语言模型通常从随机初始化开始预训练。近期研究挑战了这一惯例,表明在抽象、算法生成的数据上进行短暂预热,可以为后续的自然语言学习提供更好的起点。在本文中,我们表明,在小型语言模型中,这种预热提升了特定能力,而这些能力并未反映在语言建模困惑度中。我们的预热使用了一项需要组合能力和状态跟踪能力的抽象栈操作任务。将预训练代币中仅1%分配给该数据,即可在MUSIQUE上将多跳问答性能提升高达3.9个F1点,并在HOTPOTQA和2WIKIMULTIHOPQA上获得额外收益,尽管语言建模困惑度相当。受控实验表明,预热显著加速了更深层推理链的获取。我们还探讨了驱动这种迁移的因素。首先,数据的结构很重要:用队列任务替换栈任务无法产生相同的收益。其次,收益具有特异性:在顺序推理链上性能有所提升,但在组合或比较独立事实的任务上没有一致的收益。第三,时机很重要:将抽象数据与自然语言混合的效果远不如初始专用阶段,而在预训练之后暴露则完全消除了收益。早期优势在后续数十亿个语言代币中持续存在。这些结果表明,早期抽象训练可以可靠地塑造语言模型后续获得的能力。

英文摘要

Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract, algorithmically generated data can provide a better starting point for subsequent learning of natural language. In this paper, we show that in small language models, such a warm-up improves specific capabilities that are not reflected in language-modeling perplexity. Our warm-up uses an abstract stack-manipulation task that requires compositional and state-tracking capabilities. Allocating as little as 1% of pretraining tokens to this data improves multi-hop question answering by up to 3.9 F1 points on MUSIQUE, with additional gains on HOTPOTQA and 2WIKIMULTIHOPQA despite comparable language-modeling perplexity. Controlled experiments show that the warm-up substantially accelerates the acquisition of deeper reasoning chains. We also explore what drives this transfer. First, the structure of the data matters: replacing the stack task with a queue fails to produce the same gains. Second, the gains are specific: performance improves on sequential reasoning chains, with no consistent benefit on tasks that combine or compare independent facts. Third, timing matters: mixing abstract data with natural language is far less effective than an initial dedicated phase, and exposure after pretraining completely removes the benefits. The early advantage persists through billions of subsequent language tokens. These results show that early abstract training can reliably shape the capabilities language models later acquire.

CommentsProject page: https://zlshinnick.github.io/beyond-perplexity/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑