arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在智能手机上微调 30 亿参数大语言模型:持续训练的特征刻画

Fine-Tuning a 3B-Parameter LLM on a Smartphone: Characterizing Sustained Training

Andrew Geyko, Marius Mosbach, André Brinkmann

arXiv 2610.06325首次发表:更新:

发表机构

Saarland University(萨尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次系统刻画了在手机上微调数十亿参数大语言模型的过程,证明 iPhone 17 Pro 可在一次充电内完成 3B 模型微调,并修复了 MLX 的缺陷,使训练提速 1.47 倍、能耗降低三分之一。

AI 中文摘要

数十亿参数的大语言模型现在可以在手机上运行推理,而在设备上训练这些模型可以在不将用户数据发送到手机之外的情况下实现个性化。先前的工作测量了此类模型在手机上的单个训练步骤,但没有测量完整的训练过程,也没有测量在设备上训练的适配器是否改善个性化。我们首次对在移动设备上微调数十亿参数大语言模型进行了系统性的特征刻画,涵盖内存、每步时间、热行为和能耗。iPhone 17 Pro 可以在一次电池充电内将 30 亿参数的大语言模型微调至典型用户水平,并且生成的适配器在个性化方面的提升与在服务器上训练的适配器相当。持续训练会将手机性能限制到其初始吞吐量的大约一半,而我们测试的暂停或突发调度方案均无法恢复该性能。每个训练步骤的几乎所有时间都花在冻结的基础模型上,其中大部分时间花在反向传播中,而我们所审计的十个其他运行时中有九个不加速该过程。Apple 的 MLX 有一个内核用于此过程,但该内核从未被调度且不正确,而我们修复后的版本现已合并到上游,训练适配器速度快 1.47 倍,能耗降低三分之一。设备端微调在当前手机上可行,而使其高效需要运行时和操作系统将训练视为一等工作负载。

英文摘要

Multi-billion-parameter LLMs now run on phones for inference, and training them on the device would personalize them without user data leaving the phone. Prior work has measured individual training steps of such models on phones, but not complete training runs, and not whether adapters trained on the device improve personalization. We present the first systematic characterization of a multi-billion-parameter LLM fine-tuned on a mobile device, covering memory, per-step time, thermal behavior, and energy. An iPhone 17 Pro can fine-tune a 3B-parameter LLM to a typical user within one battery charge, and the resulting adapters improve personalization as much as adapters trained on a server. Sustained training throttles the phone to about half its initial throughput, and none of the pausing or burst schedules we tested recovers it. Nearly all of each training step is spent in the frozen base model, most of it in the backward pass, which nine of the ten other runtimes we audited do not accelerate. Apple's MLX had a kernel for it that was never dispatched and was incorrect, and our repair, now merged upstream, trains an adapter 1.47x faster on a third less energy. On-device fine-tuning is feasible on current phones, and making it efficient requires runtimes and operating systems to treat training as a first-class workload.

Comments15 pages, 7 figures, 8 tables. Code and data: https://github.com/gordofreemo/mobile_LoRA_ft

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑