arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13512cs.LGcs.AIcs.DC

自适应相位切换用于通信高效的联邦LoRA微调

A Measured Communication-Quality Frontier for Federated LoRA Fine-Tuning with Adaptive Phase-Switching

Jerry Adams Franklin

首次发表
浏览论文内容

中文总结 AI 辅助

针对联邦LoRA微调通信开销大的问题,提出自适应相位切换聚合器ReverseAdaptive,通过监测损失相对改善定位前沿拐点,在TinyLlama上实现40.5%通信节省且质量损失极小,阈值可跨规模迁移。

中文摘要 AI 辅助

联邦微调大型语言模型时,低秩适配减少了每个客户端的可训练参数,但客户端到服务器的通信仍然是主要成本。现有的联邦LoRA协议核算忽略了协议改变聚合模式时的不对称过渡轮次,并且报告的节省量忽略了分组查询注意力形状。本文测量了双向仅B端联邦LoRA协议的每轮上传和下载字节数,并将五种方法(其中三种来自先前工作)置于单一通信质量前沿上。该前沿存在一个拐点,自适应相位切换聚合器ReverseAdaptive通过监测全局训练损失的相对改善与无量纲阈值进行比较来定位该拐点,而不是预先固定相位边界。在TinyLlama-1.1B-Chat与Alpaca数据集上,ReverseAdaptive相比FLoRA实现了40.5%的实测往返节省,而保留的指令跟随损失代价为0.0063。它比FFA-LoRA(在初始化时冻结两个LoRA因子中的第一个)在保留损失上优0.0182,超过该指标上每种方法最大种子标准差二十倍以上,因此在冻结前学习该因子能产生更好的适配器。相同的阈值可跨模型规模迁移而无需重新调整,且过渡的质量代价在两个测试数据集上保持稳定。

英文摘要

Federated fine-tuning of large language models with low-rank adaptation (LoRA) reduces the number of trainable parameters, but communication remains the dominant cost, and protocols are usually compared by parameter-count ratios rather than by measured bytes. This paper measures per-round upload and download bytes for five federated LoRA protocols, three of them from prior work, and places them on a single communication-quality frontier scored by held-out instruction-following loss. The frontier has a knee. ReverseAdaptive, which learns both LoRA factors before freezing one once the relative improvement in training loss falls below a dimensionless threshold, sits at that knee: it cuts measured round-trip communication by 40.5% relative to FLoRA at a held-out loss cost of 0.0063, and beats FFA-LoRA, which freezes that factor at initialization, by 0.0182 in held-out loss, more than twenty times the largest per-method seed standard deviation. The same threshold carries to LLaMA-3.2-3B without retuning, where it saves 30.0%.

补充信息

↑