arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在无依赖栈上预训练和适配语言模型:从随机权重复现 GPT-2 124M(对照 llm.c),以及针对 Qwen3-0.6B 的临床适配器

Pretraining and adapting a language model on a dependency-free stack: GPT-2 124M from random weights, reproduced against llm.c, and a clinical adapter for Qwen3-0.6B

Thang Tran, Lan Dang

arXiv 2609.28568首次发表:更新:

发表机构

CloudKites AI Lab; Monash Business School, Monash University(CloudKites AI 实验室; 莫纳什大学莫纳什商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文用无依赖的 Zig 机器学习栈 numbat 从零预训练 GPT-2 124M,并适配小型模型至临床问答,结果与参考实现高度一致且吞吐量更优,证明了训练知识大部分属于语言模型本身而非特定软件。

AI 中文摘要

几乎每一个在用的语言模型都是由同一族软件训练出来的。这种集中性使得一个问题难以定论:关于训练语言模型的已知内容,有多少描述的是语言模型本身,又有多少描述的是那套软件?要解决这个问题,需要第二个实现,能够将模型贯穿整个生命周期,而不仅仅是复现一个算子。我们报告了这样一个生命周期。使用 numbat——一个用 Zig 编写、无第三方运行时依赖的机器学习栈——我们从随机初始化预训练了一个 124.4M 参数的 GPT-2,训练数据为 9.91B 个网页文本 token,然后将另一个小型模型适配到临床问答任务。参考实现在两个阶段均运行于相同的硬件上,一个有权中止运行的 sidecar 对每个阶段进行监督。两者结果非常接近。留出集上的交叉熵最终为 3.2588,而公开发表值为 3.29;HellaSwag 得分为 0.3053,而参考值为 0.299;在 8 次配对评估中,numbat 在每一个评估点上均低于同机参考实现,平均低 0.0608。在不同硬件上重新运行该参考实现,结果移动了 0.0035,这界定了任何差距中有多少来自方法而非框架。吞吐量并未因一致性而付出代价:在一次生产配置的单次会话中,numbat 达到了每秒 43,374 个 token,而 PyTorch 为每秒 41,202 个,在三张卡上扩展了 2.769 倍。临床适配的留出损失最终为 2.1899,而参考值为 2.1941。两个模型都不是医疗设备,也都未经验证可用于临床。

英文摘要

Almost every language model in service was trained by one family of software. That concentration makes a question hard to settle: how much of what is known about training a language model describes language models, and how much describes that software? Settling it needs a second implementation able to carry a model through a whole lifecycle rather than reproduce one operator. We report such a lifecycle. Using numbat, a machine-learning stack written in Zig with no third-party runtime dependencies, we pretrain a 124.4 M-parameter GPT-2 from random initialisation over 9.91 B tokens of web text, then adapt a separate small model to clinical question answering. A reference implementation runs on identical hardware at both stages, and a sidecar with authority to halt a run supervises each. Agreement is close. Held-out cross-entropy finishes at 3.2588 against a published 3.29, and HellaSwag at 0.3053 against 0.299; across 8 paired evaluations it sits below a same-machine reference at every point, by 0.0608 on average. Re-running that reference on different hardware moves it 0.0035, which bounds how much of any gap is method rather than framework. Throughput does not pay for agreement: measured in one session at a production configuration, numbat reaches 43,374 tokens per second against PyTorch's 41,202, scaling 2.769x over three cards. Clinical adaptation ends at 2.1899 held-out loss against 2.1941. Neither model is a medical device, and neither is validated for clinical use.

Comments20 pages, 2 figures, 7 tables. Weights: https://huggingface.co/cloudkites/gpt2-124m-fineweb-edu Licence: Apache-2.0. Weights, a config and an evaluation curve only; no framework source is released and none is needed to use them

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑