arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FLoKD:无线网络上的联邦低秩大语言模型自适应知识蒸馏

FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks

Xinlu Zhang, Na Yan, Yang Su, Yansha Deng, Toktam Mahmoodi

arXiv 2609.13580首次发表:更新:

AI 中文总结

针对无线网络中联邦大模型微调的通信瓶颈,提出FLoKD框架,通过传输中间LoRA激活并选择性传输重要块及数据集,降低50-65%通信开销,同时保持快速收敛和竞争性困惑度。

AI 中文摘要

大语言模型(LLMs)已在广泛的自然语言处理任务中展现出强大的能力。然而,传统的微调通常依赖于集中式数据收集,这带来了隐私问题。联邦学习(FL)使得在不共享原始客户端数据的情况下进行协作式LLM微调成为可能,但其在带宽受限的无线网络上的部署受到模型参数传输通信开销的阻碍。尽管低秩自适应(LoRA)减少了可训练参数的数量,但其通信成本仍随模型规模增加而增长。知识蒸馏通过输出logits避免了参数共享,但LLM中token级别的logits由于序列长度和词汇表大小而带来高通信成本。减少logits会降低成本,但会削弱监督并降低准确性。为解决这些限制,我们提出了FLoKD,一种用于无线网络上LLM联邦LoRA微调的自适应知识蒸馏框架,该框架以中间LoRA激活作为蒸馏信号进行通信,而非logits或完整参数。由于在整个公共数据集上传输所有块仍然代价高昂,我们进一步提出了一种Transformer块重要性评分框架,选择性地传输信息量最大的块,以及两种数据集选择策略,这些策略丢弃偏离本地数据分布的公共样本,并优先选择对蒸馏最有信息量的样本。在包括WikiText-103、PTB和Dialog在内的多个生成语言数据集上的大量实验表明,与基线相比,我们提出的框架将通信开销降低了50-65%,同时实现了快速收敛到具有竞争力的困惑度。

英文摘要

Large language models (LLMs) have demonstrated strong capabilities across a wide range of natural language processing tasks. However, conventional fine-tuning typically relies on centralized data collection, bringing in privacy concerns. Federated learning (FL) enables collaborative LLM fine-tuning without sharing raw client data, but its deployment over bandwidth-constrained wireless networks is hindered by the communication overhead of model-parameter transmission. Although Low-Rank Adaptation (LoRA) reduces the number of trainable parameters, its communication cost still increases with model scale. Knowledge distillation avoids parameter sharing via output logits, but token-level logits in LLMs incur high communication cost due to sequence length and vocabulary size. Reducing logits lowers the cost but weakens supervision and degrades accuracy. To address these limitations, we propose FLoKD, an adaptive knowledge-distillation framework for federated LoRA fine-tuning of LLMs over wireless networks, which communicates intermediate LoRA activations as the distillation signal rather than logits or full parameters. Since transmitting all blocks over the entire public dataset remains costly, we further propose a transformer block importance scoring framework that selectively transmits the most informative blocks, and two dataset selection strategies that discard public samples deviating from the local data distribution and prioritise those most informative for distillation. Extensive experiments across multiple generative language datasets, including WikiText-103, PTB, and Dialog, demonstrate that our proposed framework reduces communication overhead by 50-65% while achieving rapid convergence to competitive perplexity compared to baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑