面向去中心化GPU网格的通信高效LLM适配
Communication-Efficient LLM Adaptation over Decentralized GPU Meshes
浏览论文内容
中文总结 AI 辅助
针对去中心化GPU网格上的LLM预训练后适配,提出异步双回路系统与频谱校正优化器,利用锚定掩码压缩实现高达40倍以上吞吐量提升,且性能不损失。
中文摘要 AI 辅助
去中心化训练使得在低端GPU和互联网级连接上进行大规模模型训练成为可能,但数据并行和流水线并行轴上的通信成为主要瓶颈。我们研究在此设置下的预训练后适配问题。我们提出一个异步双回路系统:一个快速压缩训练回路通过激活掩码进行流水线并行(PP)传输和压缩数据并行(DP)同步来驱动吞吐量,而一个慢速锚定回路在关键路径之外运行偶尔未掩码的前向-后向传播。然后,我们引入一种频谱校正优化器,利用这些延迟的锚定先验来去噪掩码梯度,而不阻塞快速流。尽管先前的工作发现激进的激活压缩不可靠,但我们表明,在这种锚定方式下,掩码支持高压缩率下的预训练后适配。仅流水线并行压缩即可带来高达$9\ imes$的吞吐量提升,将其与数据并行压缩结合,在互联网级约200Mbps连接上提升超过$40\ imes$,同时在与密集未压缩性能匹配的情况下,跨越领域适配和持续预训练。
英文摘要
Decentralized training enables large-model training over low-end GPUs and internet-grade connections, but communication along both data-parallel and pipeline-parallel axes becomes the primary bottleneck. We study post-pretraining adaptation in this setting. We propose an asynchronous two-circuit system: a fast compressed training circuit drives throughput using activation masking for pipeline-parallel (PP) transfer and compressed data-parallel (DP) synchronization, while a slow anchor circuit runs occasional unmasked forward--backward passes off the critical path. Then, we introduce a spectral correction optimizer that uses these delayed anchor priors to denoise masked gradients without blocking the fast stream. Although prior work has found aggressive activation compression unreliable, we show that masking supports post-pretraining adaptation at high compression rates when anchored this way. Pipeline-parallel compression alone yields up to a $9\times$ throughput gain, and combining it with data-parallel compression increases beyond $40\times$ over internet-grade $\sim 200$Mbps connections, while matching dense uncompressed performance across domain adaptation and continual pretraining.
发表机构
- Pluralis Research
机构由 AI 辅助整理,请以论文原文为准。