arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoeF-SFL:保留协作式客户端-服务器学习并增强通信效率

CoeF-SFL: Preserving Collaborative Server-Client Learning with Enhanced Communication Efficiency

Junwoo Bae, Jin-Hyun Ahn

arXiv 2609.34360首次发表:更新:

发表机构

Myongji University(明知大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对分割联邦学习通信开销大且辅助网络破坏端到端目标的问题,提出CoeF-SFL,通过每轮一次交换和曲率校正补偿,显著提升通信效率与协作性能。

AI 中文摘要

分割联邦学习(SFL)使资源受限的客户端能够参与协作训练,但原始SFL在每个批次交换粉碎数据和梯度,这导致显著的通信开销。近期方法通过在客户端切割层引入辅助网络来减少此开销。然而,我们发现这种方法使客户端优化的局部目标与端到端目标不同,从根本上限制了客户端与服务器之间的协作训练。我们提出基于补偿反馈的SFL(CoeF-SFL),一种无需任何辅助网络且保留端到端目标的通信高效框架。在CoeF-SFL中,客户端和服务器每轮交换一次粉碎数据和梯度,并在本地训练期间重复使用它们。由于这种重复使用使客户端侧的梯度变得过时,我们在激活空间中用基于曲率的校正补偿它们,并开发了两种变体。CoeF-D用对角梯度外积近似Hessian,而CoeF-J利用一个上界真实损失的替代损失的可处理的基于Jacobian的Hessian。我们提供了每种方法的理论背景,刻画了其补偿。在视觉和语言任务、模型容量、切割层和数据分布方面,CoeF-SFL在相同通信频率下显著优于基于辅助网络的方法,且在视觉任务上改进最为显著。代码可在该https URL获取。

英文摘要

Split Federated Learning (SFL) enables resource-constrained clients to participate in collaborative training, but vanilla SFL exchanges smashed data and gradients at every batch, which incurs significant communication overhead. Recent methods reduce this overhead with an auxiliary network at the client-side cut layer. However, we identify that this approach makes the client optimize a local objective that differs from the end-to-end objective, which fundamentally limits the collaborative training between the client and the server. We propose Compensated Feedback based SFL (CoeF-SFL), a communication-efficient framework that retains the end-to-end objective without any auxiliary network. In CoeF-SFL, the client and the server exchange the smashed data and the gradients once per round and reuse them during local training. Since this reuse makes the gradients stale on the client side, we compensate them with a curvature-based correction in the activation space and develop two variants. CoeF-D approximates the Hessian with a diagonal gradient outer product, while CoeF-J exploits the tractable Jacobian-based Hessian of a surrogate loss that upper-bounds the true loss. We provide the theoretical background of each method, characterizing its compensation. Across vision and language tasks, model capacities, cut layers, and data distributions, CoeF-SFL significantly outperforms auxiliary-network-based methods under the same communication frequency, and the improvement is most substantial on vision tasks. Code is available at https://anonymous.4open.science/r/CoeF-SFL-2686/README.md

CommentsSubmitted to a conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑