迈向私有大语言模型训练:探索在Apple Silicon上通过Thunderbolt RDMA进行语言模型微调
Towards Training Private LLMs: Exploring Fine-Tuning Language Models on Apple Silicon with RDMA over Thunderbolt
浏览论文内容
中文总结 AI 辅助
本文探索Apple Silicon作为私有LLM微调平台的可行性,通过多链路通信、持久线程和梯度重叠优化,在四节点Mac Studio集群上微调Qwen3-9B,吞吐量提升1.6倍,并证明其成本效益优于H100。
中文摘要 AI 辅助
私有大语言模型(LLM)微调对于需要使用敏感数据调整模型的组织而言日益重要,但这通常超出商用数据中心加速器的内存容量。Apple Silicon通过大统一内存和较低的系统总成本提供了不同的设计点,而近期Apple软件支持使得通过Thunderbolt上的RDMA(RDMA-over-Thunderbolt,TB)进行分布式执行成为可能。本文研究Apple Silicon能否作为私有LLM微调的实用平台。我们表征了Mac Studio节点上RDMA-over-TB通信,显示实测带宽远低于TB标称规格。接下来,我们扩展了Apple的实现,加入多链路通信、持久工作线程和CPU端梯度重叠,以更好地利用多个直接TB链路用于LLM微调工作负载。最后,在一个微调Qwen3-9B的四节点Mac Studio集群上,我们的优化将弱扩展吞吐量相比单链路、无重叠基线提升高达1.6倍,并在序列长度17408时达到936 tokens/s。我们进一步将Apple Silicon与NVIDIA H100平台进行比较,以量化内存容量、吞吐量和购置成本之间的权衡,表明Apple Silicon可为私有LLM微调提供成本有效的解决方案。
英文摘要
Private large language model (LLM) fine-tuning is increasingly important for organizations that need to adapt models using sensitive data, but it often exceeds the memory capacity of commodity datacenter accelerators. Apple Silicon offers a different design point through large unified memory and lower complete-system cost, while recent Apple software support enables distributed execution over RDMA-over-Thunderbolt (TB). This paper studies whether Apple Silicon can serve as a practical platform for private LLM fine-tuning. We characterize RDMA-over-TB communication on Mac Studio nodes, showing that the measured bandwidth is far below nominal TB specifications. Next, we extend Apple's implementation with multi-trunk communication, persistent worker threads, and CPU-side gradient overlap to better exploit multiple direct TB links for LLM fine-tuning workloads. Finally, on a four-node Mac Studio cluster that fine-tunes a Qwen3-9B, our optimizations improve weak-scaling throughput by up to 1.6X over the single-trunk, non-overlapped baseline and reach 936 tokens/s for sequence length 17408. We further compare Apple Silicon with an NVIDIA H100 platform to quantify the trade-off between memory capacity, throughput, and acquisition cost, showing that Apple Silicon can provide a cost-effective solution for private LLM fine-tuning.
发表机构
- National Taiwan University(台湾大学)
- Academia Sinica(中央研究院)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。