arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Xronos:面向边缘 CPU 协作式大语言模型微调的异构感知张量并行

Xronos: Heterogeneity-Aware Tensor Parallelism for Collaborative LLM Fine-Tuning on Edge CPUs

Wonmi Choi, Sunjae Park, Dohyeok Kwon, Zhixiong Niu, Yeonho Yoo, Chuck Yoo, Gyeongsik Yang

arXiv 2609.19909首次发表:更新:

发表机构

Korea University; Microsoft Research Asia; Dongguk University(韩国大学; 微软亚洲研究院; 东国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对 CPU 边缘设备上协作式微调中流水线并行导致的 CPU 争用和异构张量并行的同步等待问题,提出 Xronos 框架,采用异构感知张量划分,显著减少微调时间和设备空闲时间。

AI 中文摘要

边缘设备上的协作式微调使大型语言模型能够适应特定领域的数据,同时保持每个设备的数据本地化。最先进的协作式微调技术主要针对基于 GPU 的边缘设备设计,并依赖流水线并行。然而,许多边缘平台(包括物联网网关、智能家居中枢和车载计算机)主要基于 CPU。本文报告指出,流水线并行在基于 CPU 的边缘设备上效率低下,因为同一 CPU 同时处理模型计算和通信,导致严重的 CPU 争用。我们的分析表明,这导致计算停顿比率平均比 GPU 设备高 5.75 倍。张量并行可以通过分离计算和通信来缓解这种争用,但现有的张量并行技术假设设备是同构的。在异构 CPU 边缘设备上,我们发现这一假设导致较快的 worker 在同步点等待较慢设备时,空闲时间高达 34%。为解决这些局限,我们提出了 Xronos,一个面向异构 CPU 边缘设备的协作式微调框架。Xronos 以张量并行作为执行主干,并结合轻量级剖析与异构感知的张量划分来减少掉队者瓶颈。在多种设备、模型和基准任务上,Xronos 将微调时间减少了 18%(相对于张量并行)至 56%(相对于流水线并行),并将设备空闲时间比率相对于最先进技术降低了约 5.9 倍,同时保持了准确性。

英文摘要

Collaborative fine-tuning on edge devices adapts large language models to domain-specific data while keeping each device's data local. State-of-the-art (SOTA) collaborative fine-tuning techniques are largely designed for GPU-based edge devices and rely on pipeline parallelism (PP). However, many edge platforms, including IoT gateways, smart-home hubs, and in-vehicle computers, are primarily CPU-based. This paper reports that PP is ineffective on CPU-based edge devices because the same CPU handles both model computation and communication, which causes severe CPU contention. Our analysis shows that this leads to 5.75$\times$ higher computation stall ratios than on GPU devices on average. Tensor parallelism (TP) can alleviate this contention by separating computation and communication, but existing TP techniques assume homogeneous devices. On heterogeneous CPU edge devices, we find that this assumption causes faster workers to remain idle for up to 34% while waiting for slower devices at synchronization points. To address the limitations, we propose Xronos, a collaborative fine-tuning framework for heterogeneous CPU edge devices. Xronos uses TP as its execution backbone and combines lightweight profiling with heterogeneity-aware tensor partitioning to reduce the straggler bottleneck. Across diverse devices, models, and benchmark tasks, Xronos reduces fine-tuning time by 18% (TP) to 56% (PP) and the ratio of device idle time by $\sim$5.9$\times$ over SOTA techniques, while maintaining the accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑