arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03930cs.CLcs.AIcs.LG

语言之前的逻辑:基于形式推导的预预训练可促进技能获取与可压缩性

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

Jo-Ku Cheng, Nikolaos Aletras, Marco Valentino

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出Logic-PPT策略,通过形式推导预预训练语言模型,可加速技能获取、提升性能,还能增强模型可压缩性,在少用36B token时达80%准确率,33%稀疏度下匹配密集基线性能。

中文摘要 AI 辅助

在符号数据上对语言模型(LMs)进行预预训练能够加速并提升自然语言获取效果,但现有预预训练任务(如Dyck和过程算法)依赖狭窄的原语,无法捕捉自然语言的表达能力;且现有研究局限于相对较小的token预算,对技能涌现和表征动态的见解有限。为解决这些局限,我们提出逻辑预预训练(Logic-PPT)作为一种有原则的初始化策略,利用形式推导赋予更丰富的结构和语言偏见。形式推导需要抽象机制,这些机制是自然语言的核心,可同时绑定变量、连接量词与关系依赖,并在长上下文上组合谓词-论元结构。我们将评估规模扩展至100B token级别,结果显示逻辑预预训练大幅加速了语言模型的技能获取:在语言任务上达到80%准确率,比标准初始化少用36B token,且优于其他预预训练基线。从机制上看,形式推导诱导出持续的结构重组,其特征是表征空间的秩更低、频谱更集中;关键的是,我们证明这种内部几何结构可通过剪枝提升模型可压缩性,在约33%稀疏度下仍能匹配密集基线的性能。

英文摘要

Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. Moreover, prior studies remain restricted to relatively small token budgets, offering limited insight into skill emergence and representational dynamics. To address these limitations, we propose logic pre-pretraining (Logic-PPT) as a principled initialization strategy, leveraging formal derivations to impart richer structural and linguistic biases. Formal derivations require abstract mechanisms that are central to natural language, simultaneously binding variables, connecting quantifiers and relational dependencies, and composing predicate-argument structures over long contexts. Scaling our evaluation to a 100B-token regime, logic pre-pretraining substantially accelerates skill acquisition in LMs, achieving 80\% accuracy on linguistic tasks with 36B fewer tokens than standard initialization, and outperforming alternative pre-pretraining baselines. Mechanistically, formal derivations induce persistent structural reorganization, distinctively characterized by a lower-rank, spectrally concentrated representation space. Crucially, we show that this internal geometry enables improved model compressibility via pruning, matching the dense baseline performance even at $\approx$33\% sparsity.

↑