ZeroLock:通过模块化更新解耦实现并发内存高效的大语言模型训练
ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling
浏览论文内容
中文总结 AI 辅助
针对边缘LLM微调的BP训练存在更新锁定瓶颈,本研究提出ZeroLock无BP算法,通过模块化更新解耦突破局限,经实验验证可降内存26.5%、提吞吐量4.9%。
中文摘要 AI 辅助
边缘设备上的大语言模型(LLM)微调可使模型适配特定场景数据,同时保护隐私。尽管现有研究提出流水线并行性来解决边缘设备内存与计算资源有限的问题,但这些方法通常依赖反向传播(BP)训练,而BP存在更新锁定的根本局限,可能出现严重的吞吐量和内存瓶颈。本研究提出一种名为ZeroLock的无BP算法,通过构建局部目标将模型更新解耦为独立的分块更新,打破BP的更新锁定,从而在算法层面提升吞吐量,并通过减少激活存储降低内存使用。据我们所知,我们首次为这类基于局部目标构建的方法提供了通用模型分块划分下的理论框架,将局部目标映射至全局目标。我们证明ZeroLock的收敛速率为$\tilde{\text{O}}(1/\text{sqrt}(T))$,与BP的收敛速率仅相差多对数因子。我们设计了ZeroLock的系统并构建了实际原型,融入了提前转发和故障恢复等技术以实现高效且鲁棒的部署。原型实验结果显示,与基于BP的基线相比,ZeroLock将内存降低26.5%,吞吐量提升4.9%。
英文摘要
Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy. Although existing studies proposed pipeline parallelism to address the limited memory and computing resources of edge devices, they commonly rely on backpropagation (BP) training, which has a fundamental limitation of update locking and could experience severe throughput and memory bottlenecks. In this work, we propose a BP-free algorithm, called ZeroLock, that decouples the model updates into independent chunk updates by local objective construction. It breaks the update locking of BP and hence can improve throughput at the algorithm level and lower memory usage by reducing activation storage. To the best of our knowledge, we provide the first theoretical framework for such local objective construction-based approaches under general model chunk division by mapping local objectives to the global objective. We prove that ZeroLock has a convergence rate of $\tilde{\mathcal{O}}(1/\sqrt{T})$, which differs from BP only by polylogarithmic factors. We design a system for ZeroLock and build real-world prototypes, incorporating techniques such as early forwarding and failure recovery for efficient and robust implementation. Experiments on the prototype show that compared to BP-based baselines, ZeroLock reduces the memory by 26.5% and improves throughput by 4.9%.
发表机构
- Southern University of Science and Technology(南方科技大学)
- Montclair State University(蒙特克莱尔州立大学)
机构由 AI 辅助整理,请以论文原文为准。