arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18066cs.DC

迈向私有大语言模型训练:探索在Apple Silicon上通过Thunderbolt RDMA进行语言模型微调

Towards Training Private LLMs: Exploring Fine-Tuning Language Models on Apple Silicon with RDMA over Thunderbolt

En-Ming Huang, Yao-Ting Hsieh, Hsiang-Yu Tsou, Mu-Chi Chen, Shih-Hao Hung, H. T. Kung

首次发表
浏览论文内容

中文总结 AI 辅助

本文探索Apple Silicon作为私有LLM微调平台的可行性,通过多链路通信、持久线程和梯度重叠优化,在四节点Mac Studio集群上微调Qwen3-9B,吞吐量提升1.6倍,并证明其成本效益优于H100。

中文摘要 AI 辅助

私有大语言模型(LLM)微调对于需要使用敏感数据调整模型的组织而言日益重要,但这通常超出商用数据中心加速器的内存容量。Apple Silicon通过大统一内存和较低的系统总成本提供了不同的设计点,而近期Apple软件支持使得通过Thunderbolt上的RDMA(RDMA-over-Thunderbolt,TB)进行分布式执行成为可能。本文研究Apple Silicon能否作为私有LLM微调的实用平台。我们表征了Mac Studio节点上RDMA-over-TB通信,显示实测带宽远低于TB标称规格。接下来,我们扩展了Apple的实现,加入多链路通信、持久工作线程和CPU端梯度重叠,以更好地利用多个直接TB链路用于LLM微调工作负载。最后,在一个微调Qwen3-9B的四节点Mac Studio集群上,我们的优化将弱扩展吞吐量相比单链路、无重叠基线提升高达1.6倍,并在序列长度17408时达到936 tokens/s。我们进一步将Apple Silicon与NVIDIA H100平台进行比较,以量化内存容量、吞吐量和购置成本之间的权衡,表明Apple Silicon可为私有LLM微调提供成本有效的解决方案。

英文摘要

Private large language model (LLM) fine-tuning is increasingly important for organizations that need to adapt models using sensitive data, but it often exceeds the memory capacity of commodity datacenter accelerators. Apple Silicon offers a different design point through large unified memory and lower complete-system cost, while recent Apple software support enables distributed execution over RDMA-over-Thunderbolt (TB). This paper studies whether Apple Silicon can serve as a practical platform for private LLM fine-tuning. We characterize RDMA-over-TB communication on Mac Studio nodes, showing that the measured bandwidth is far below nominal TB specifications. Next, we extend Apple's implementation with multi-trunk communication, persistent worker threads, and CPU-side gradient overlap to better exploit multiple direct TB links for LLM fine-tuning workloads. Finally, on a four-node Mac Studio cluster that fine-tunes a Qwen3-9B, our optimizations improve weak-scaling throughput by up to 1.6X over the single-trunk, non-overlapped baseline and reach 936 tokens/s for sequence length 17408. We further compare Apple Silicon with an NVIDIA H100 platform to quantify the trade-off between memory capacity, throughput, and acquisition cost, showing that Apple Silicon can provide a cost-effective solution for private LLM fine-tuning.

发表机构

  • National Taiwan University(台湾大学)
  • Academia Sinica(中央研究院)
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑