发表机构
Fraunhofer Institute for Algorithms and Scientific Computing SCAI; University Hospital Bonn(弗劳恩霍夫算法与科学计算研究所; 波恩大学医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出先验拟合语言模型(PFLM),一种仅基于合成非语言先验训练的300M参数字节级Transformer,通过推断上下文语言实现预测,并在多语言维基百科及非文本领域上展现出强大的压缩与学习能力。
AI 中文摘要
我们提出了先验拟合语言模型(PFLM),这是一个300M参数的字节级Transformer,仅在来自合成非语言先验的样本上进行预训练。给定真实文本的前缀,它能在冻结权重的情况下学习预测上下文中的语言,而从未见过任何真实语言的单词。每个训练序列都是由一个从这类模型分布中新鲜抽取的循环结构因果模型生成的。在训练过程中,模型从未两次看到相同的语言,因此预测续写的唯一方法是从前缀中推断出语言。来自该先验的样本共享自然文本的统计特征:齐普夫频率、缓慢的熵率收敛和长程依赖性。在六种语言的维基百科上,每字节比特数从均匀的8降至0.9到2.4之间,在上下文达到一百万字节时。给定数字而非文本,PFLM学会计数、比较大小和近似相加。它能预测确定性序列,如鲁丁-夏皮罗序列或素数指示符,并且能将六个非文本领域(从源代码到语音)压缩到低于gzip和PPMd的水平。该模型并未学会一种语言,而是学会了如何学习一种语言。
英文摘要
We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only way to predict the continuation is to infer the language from the prefix. Samples from this prior share the statistical signatures of natural text: Zipfian frequencies, slow entropy-rate convergence, and long-range dependence. On Wikipedia in six languages, bits per byte fall from the uniform eight to between 0.9 and 2.4 at one million bytes of context. Given numerals instead of text, PFLM learns to count, to compare magnitudes, and to add approximately. It predicts deterministic sequences like Rudin-Shapiro or the prime indicator, and it compresses six non-text domains, from source code to speech, below gzip and PPMd. The model has not learned a language. It has learned to learn one.
Comments15 pages, 6 figures, 5 tables, Code: https://github.com/cbl/prior-fitted-language-model, weights: https://huggingface.co/lennartcb/pflm1