发表机构
Vellore Institute of Technology; University of Arizona(韦洛尔理工学院; 亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对拆分式联邦微调(SFF),识别出深度-性能困境,评估多种联邦适配器聚合方法后发现其无法缓解拆分架构缺陷,归因于Transformer拓扑,为分布式LLM微调提供结构基础。
AI 中文摘要
拆分式联邦微调(Split Federated Fine-tuning, SFF)是一种通过在资源受限的客户端与中央服务器之间划分模型深度来扩展大语言模型(Large Language Models, LLMs)的有前景范式。尽管吞吐量和隐私的系统激励倾向于深度划分,但此类配置对模型效用的影响却鲜为人知。在本研究中,我们识别并表征了深度-性能困境:最大化系统效率的机制恰好是微调质量崩溃的区域。通过对四种模型规模(从GPT-2到Llama-3-8B)及多种基准的全面审计,我们证明更深的划分会带来吞吐量和隐私的单调提升,但代价是灾难性的性能停滞。我们评估了一系列最先进的联邦适配器聚合方法,包括AVG、STACK、SVD和FREEZE,发现这些技术在标准联邦学习中虽有效,但无法缓解拆分架构特有的缺陷。最后,我们对这种失败进行了机制诊断,将崩溃追溯至Transformer的近等距拓扑结构,该结构使聚合噪声在衰减前传播,直至触发服务器划分中的注意力崩溃。我们的发现挑战了划分深度是效用中性调谐旋钮的普遍假设,并为稳定的分布式LLM微调提供了结构基础。
英文摘要
Split Federated Fine-tuning (SFF) is a promising paradigm for scaling Large Language Models (LLMs) by partitioning model depth between resource-constrained clients and a centralized server. While system incentives for throughput and privacy favor deep partitions, the impact of such configurations on model utility remains poorly understood. In this work, we identify and characterize the Depth-Performance Dilemma: the regime that maximizes system efficiency is precisely where fine-tuning quality collapses. Through a comprehensive audit across four model scales (GPT-2 to Llama-3-8B) and diverse benchmarks, we demonstrate that deeper partitions provide monotonic gains in throughput and privacy at the cost of catastrophic performance plateaus. We evaluate a suite of state-of-the-art federated adapter aggregation methods including AVG, STACK, SVD, and FREEZE, revealing that while these techniques are effective in standard Federated Learning, they fail to mitigate the artifacts unique to split architectures. Finally, we provide a mechanistic diagnosis for this failure, tracing the collapse to the near-isometric topology of Transformers, which allows aggregation noise to propagate without attenuation until it triggers Attention Collapse in the server partition. Our findings challenge the prevailing assumption that partition depth is a utility-neutral tuning knob and provide a structural foundation for stable distributed LLM fine-tuning.
Comments24 pages, 10 figures, 10 tables