arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型的预预训练:并非总能带来帮助——多语言调查

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

Sofiia Riazhskykh, Nam Luu, Ondřej Bojar

arXiv 2608.08800首次发表:更新:

发表机构

Charles University; Faculty of Mathematics and Physics(查理大学; 数学与物理学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究调查了 LLM 预预训练在多语言场景下的 token 效率增益,发现其受实验设置和随机种子影响,仅在特定条件下稳定,建议部分实验需多次运行以规避不稳定方法。

AI 中文摘要

对大语言模型(LLM)进行人工语言预训练(即“预预训练”)是一项据称可将 token 效率提高 33% 的技术,即最多可节省达到特定性能所需的 33% 训练 token。我们针对英语在四个语系的更多自然语言上验证了这一先前结果,使用了两种不同的 tokenizer 并改变了模型规模。我们还将观察到的 token 效率增益(或损失)与语言的量化语言特性相关联,例如句子长度、形态丰富度以及依存句法树的特征(树深度、子节点数量、交叉依存数量)。我们的实证结果表明,所报告的增益在很大程度上取决于实验设置和随机种子的选择,不过我们可以确认,对于大多数受检语言,使用 Llama tokenizer 对小型模型进行 128-Dyck 预训练时,增益趋势是稳定的。总体而言,我们认为至少应对部分实验进行多次训练运行,以避免学术界采用不稳定的方法。

英文摘要

Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33% of training tokens needed to reach a certain performance. We validate this prior result for English on a larger set of natural languages across four language families, using two different tokenizers and varying model sizes. We also relate the observed gains (or losses) in token efficiency to quantified linguistic properties of the languages, such as sentence length, morphological richness, and features of dependency syntactic trees (tree depth, number of children, number of crossing dependencies). Our empirical results indicate that the reported gains depend heavily on the experiment setup and the choice of random seed, although we can confirm the trend of stable gains with 128-Dyck pretraining of small models with the Llama tokenizer for most of the examined languages. On a general note, we argue that multiple training runs should be carried out at least for a subset of experiments to avoid the community adopting unstable approaches.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑