Agora:大规模语言模型的集体无许可互联网规模预训练
Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models
浏览论文内容
中文总结 AI 辅助
研究旨在解决大语言模型训练资源受限问题,提出Agora系统,通过互联网级链路的带宽高效流水线并行模型分片与多方容错集体操作相结合,实现集体训练、集体所有模型,首次展示Pluralis - 8B预训练,效率达集中式基线63%。
中文摘要 AI 辅助
训练数十亿到数万亿参数规模的大语言模型局限于数据中心,数据并行(DP)和模型并行(MP)技术假定有同质加速器、高速互连和单个编排实体。前沿模型开发集中在少数能组建此类集群的团队。大量计算资源因异构、可抢占、个体所有且仅通过互联网连接而无法用于训练。我们提出Agora系统,它通过互联网级链路高效利用带宽的流水线并行模型分片与多方容错集体操作相结合。每个参与者仅持有模型的一个阶段,无单一方拥有完整权重。我们称此设置为协议学习,它实现集体训练、集体所有的模型,为具有经济可持续性的开源前沿训练开辟道路。本报告展示了在通信高效并行、异步优化和容错系统设计方面的研究成果。首次展示了Pluralis - 8B,一个在FineWeb - Edu的500B令牌上对86亿参数模型进行的开放无许可预训练。该模型由330个贡献者节点在40天内训练而成,主要是通过互联网连接的消费级GPU,节点全程可加入和离开。运行维持约170k令牌/秒和每TFLOP池化计算4.2个令牌,为集中式H100基线效率的63%,并收敛到与集中式参考运行相差很小的范围内。
英文摘要
Training large language models at the multi-billion to trillion parameter scale is confined to datacenters, where data-parallel (DP) and model-parallel (MP) techniques presume homogeneous accelerators, high-speed interconnects, and a single orchestrating entity. Frontier model development is thereby concentrated among the few groups able to assemble such clusters. Meanwhile, an enormous pool of compute remains unusable for training: consumer and professional GPUs that are heterogeneous, preemptible, individually owned, and connected only by the internet. We present Agora, a system that makes efficient use of this compute. Agora combines bandwidth-efficient pipeline-parallel model sharding over internet-grade links with multi-party, fault-tolerant collective operations. Each participant holds only one stage of the model, and no single party ever possesses the full weights. We term this setup Protocol Learning: it enables collectively trained, collectively owned models, opening a path to open-source frontier training with economic sustainability. This report presents the outcome of a research effort spanning communication-efficient parallelism, asynchronous optimization, and fault-tolerant systems design. It culminates in the first demonstration of its kind: Pluralis-8B, an open, permissionless pretraining run of an 8.6B-parameter model on 500B tokens of FineWeb-Edu. The model was trained over 40 days by 330 contributor nodes, predominantly consumer GPUs on internet connections, joining and leaving throughout. The run sustained ~170k tokens/s and 4.2 tokens per TFLOP of pooled compute, 63% of the efficiency of a centralized H100 baseline, and converged to within a small margin of a centralized reference run.