arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CommuteProp:面向通信受限的分裂式大语言模型微调的分离式训练

CommuteProp: Decoupled Training for Communication Bound Split LLM Fine-Tuning

CHEN Ding, LUO Haochen, LIU Chen

arXiv 2610.05105首次发表:更新:

发表机构

City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对分裂式大语言模型微调中的通信-计算瓶颈,提出异步解耦算法CommuteProp,重叠计算与通信以缩短周期时间,并提升吞吐量且保持精度。

AI 中文摘要

分裂学习已成为隐私保护的大语言模型微调的一种有前景的范式,然而其实际部署受到顺序通信-计算瓶颈的严重阻碍。在传统的同步流水线中,客户端在等待服务器端梯度时处于空闲状态,导致训练效率大幅下降。我们提出了CommuteProp,一种异步分裂学习算法,它将训练过程解耦为两个并发的阶段:跨块前向-反向传播和块内权重更新。通过将计算与通信重叠,CommuteProp将边际周期时间从所有阶段延迟的总和减少到主导的计算瓶颈。我们提供了严格的异步误差和收敛性分析。此外,基于我们的分析,我们推导出一种NS预条件子方法,以进一步减轻由陈旧性引起的噪声。综合实验表明,我们的算法在保持与同步方法相当的准确性的同时,实现了显著的吞吐量提升;此外,它作为一个即插即用模块,不仅适应而且积极增强了现有的隐私增强方法。

英文摘要

Split learning has emerged as a promising paradigm for privacy-preserving LLM fine-tuning, yet its practical deployment is severely hindered by the sequential communication-computation bottleneck. In conventional synchronous pipelines, clients remain idle while waiting for server-side gradients, resulting in substantial training inefficiency. We propose CommuteProp, an asynchronous split-learning algorithm that decouples the training process into two concurrent phases: a cross-block forward-backward pass and an in-block weight update. By overlapping computation with communication, CommuteProp reduces the marginal cycle time from a sum of all stage latencies to the dominant computational bottleneck. We provide a rigorous asynchronous error and convergence analysis. Moreover, we derive an NS preconditioner method based on our analysis to further mitigate staleness-induced noise. Comprehensive experiments indicate that our algorithm achieves substantial throughput gains while maintaining accuracy comparable to synchronous methods; furthermore, it functions as a plug-and-play module that not only accommodates but actively enhances existing privacy enhancement methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑