arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Puro-2B:Poor Lab基于RTX 5090训练的Qwen2-1.5B模型,训练成本低于5090美元

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

Kairong Luo, Jiarui Cui, Yaorui Yin, Shengqi Chen, Yiming Yang, Linxiang Gao, Yanmohan Wang, Chengxia Li, Mingzhe Zhang, Kaifeng Lyu, Wenguang Chen

arXiv 2608.27370首次发表:更新:

发表机构

Tsinghua University; Pengcheng Laboratory(清华大学; 鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出开源预训练方案,在RTX 5090上低成本训练出Puro-2B模型,推导出成本缩放定律并开展数据课程影响下游性能的案例研究,完整方案已开源。

AI 中文摘要

语言模型预训练几乎已与高昂成本划等号,这使得学术和开源社区的大部分群体难以参与其中。尽管目前已有包括开放权重模型和开源训练方案在内的强大开源成果,但始终缺乏一种兼具成本效益、硬件可及性与开源属性的预训练方案。即便在小规模下,训练Llama-3.2-3B的成本也超过150万美元,复现SmolLM3-3B则需要70万美元以上。本报告提出一种旨在降低该门槛的开源预训练方案,借助该方案,我们在消费级RTX 5090 GPU上以FP8精度从零开始训练了一系列Puro-2B模型,训练token量最高达1.4万亿,该系列模型在token预算和所选方案变体上存在差异。我们的最优模型训练成本不足6900美元,在我们的评估协议下其性能接近Qwen2.5-1.5B。这种成本效益源于硬件选择、低精度训练、超球优化、课程模型平均及数据方案等多种方法的结合。除方案本身外,我们还提供两项额外成果:其一,从Puro-2B系列中推导出Puro成本缩放定律,该定律关联训练成本与平均模型性能,拟合结果显示约4400美元(低于5090美元)即可达到Qwen2-1.5B的性能;其二,作为端到端案例研究,我们探究预训练数据课程如何影响后训练后的下游性能,只有拥有完整预训练流程而非仅模型权重,才能开展此类受控研究。我们依据Apache 2.0协议在该httpsURL发布了Puro-2B的完整训练方案,包含数据、代码及模型权重。

英文摘要

Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing. Even at a small scale, training Llama-3.2-3B costs over \$1.5M, and reproducing SmolLM3-3B needs over \$700K. In this report, we present an open pretraining recipe designed to lower this barrier. Using this recipe, we train a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs. The models in the collection differ in token budgets and selected recipe variants. Our best model is trained at a compute cost of less than \$6.9K and approaches Qwen2.5-1.5B performance under our evaluation protocol. This cost efficiency is enabled by a combination of approaches, including hardware selection, low-precision training, hyperball optimization, curriculum model averaging, and the data recipe. Beyond the recipe itself, we provide two additional results. First, across the Puro-2B collection, we derive a Puro Cost Scaling Law that relates training cost to average model performance; the fitted law suggests that about \$4.4K, less than \$5,090, is sufficient to reach the performance of Qwen2-1.5B. Second, as an end-to-end case study, we examine how pretraining data curricula shape downstream performance after post-training. Such controlled studies are enabled by having access to the full pretraining pipeline rather than model weights alone. We release the full training recipe for Puro-2B, including data, code, and model weights under Apache 2.0 at https://huggingface.co/collections/thu-pacman/puro-2b.

Comments63 pages, 20 figures, 24 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑