何时足够浅?具有客户端特定充分性估计的自适应拆分联邦学习
When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation
- The University of Hong Kong(香港大学)
- Sun Yat-sen University(中山大学)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对拆分联邦学习中静态拆分策略无法适配客户端异质性的问题,提出FedSGA框架,通过客户端特定充分性估计实现自适应拆分,提升模型性能并降低客户端计算量。
AI中文摘要:
拆分联邦学习(Split Federated Learning, SFL)通过在服务器与客户端之间拆分网络,实现分布式模型训练。然而在客户端异质性场景下,传统静态拆分策略可能并非最优,因为客户端在数据分布、适应动态性及表征学习进展上存在差异,单一拆分点不足以适配客户端特定的训练状态。本文提出FedSGA框架,这是一种基于充分性引导的自适应拆分联邦学习框架,通过客户端特定的浅层充分性估计解决上述问题。首先,我们引入基于私有提示令牌的客户端特定适应通道,该通道与共享骨干网络分离,单独跟踪本地适应动态性,为检测客户端适应是否仍处于活跃状态提供轻量信号。为进一步避免在多个候选深度上重复在线探测,我们设计了浅层充分性估计器,其结合跨客户端语义对齐、时间接口稳定性及提示状态变化,以估计最浅拆分是否已足够。最后,我们引入与拆分兼容的接口协调模块,将不同拆分深度的激活投影到共享语义空间,提升服务器端预测前异构客户端接口的可比性。在多个异质性基准上开展的大量实验表明,与现有最优方法相比,FedSGA在提升模型性能的同时,减少了不必要的客户端侧计算。
英文摘要:
\textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, under client heterogeneity, the conventional static split strategy may be suboptimal because clients can differ in data distributions, adaptation dynamics, and representation learning progress, making a single split point insufficient to accommodate client-specific training states. In this paper, we propose \textsc{FedSGA}, a \textbf{S}ufficiency-\textbf{G}uided \textbf{A}daptive split \textbf{Fed}erated learning framework that addresses this question through client-specific shallow sufficiency estimation. First, we introduce a client-specific adaptation channel based on private prompt tokens, which tracks local adaptation dynamics separately from the shared backbone and provides a lightweight signal for detecting whether client adaptation remains active. To further avoid repeated online probing over multiple candidate depths, we design a shallow sufficiency estimator that combines cross-client semantic alignment, temporal interface stability, and prompt-state variation to estimate whether the shallowest split is already sufficient. Finally, we introduce a split-compatible interface harmonization module that projects activations from different split depths into a shared semantic space, improving the comparability of heterogeneous client interfaces before server-side prediction. Extensive experiments on multiple heterogeneous benchmarks demonstrate the effectiveness of \textsc{FedSGA} in improving model performance compared with state-of-the-art methods while reducing unnecessary client-side computation.