arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边缘网络中联邦学习工作流管理的自适应客户端聚类与协调

Adaptive Client Clustering and Coordination for Federated Learning Workflow Management in Edge Networks

Jieping Luo, Qiyue Li, Yuxuan Chen, Hang Qi, Jiaying Yin, Jingjin Wu, Qian Wang

arXiv 2609.33544首次发表:更新:

发表机构

University of Oxford; Beijing Normal-Hong Kong Baptist University; The First Affiliated Hospital, Sun Yat-Sen University; Zhejiang University of Technology(牛津大学; 北京师范大学-香港浸会大学联合国际学院; 中山大学附属第一医院; 浙江工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对边缘网络中依赖型联邦学习工作流,提出A-CoDa框架,通过LDD聚类、FedMIX参与机制和DAG调度器,在保证精度的同时降低端到端完成时间。

AI 中文摘要

联邦学习(FL)越来越多地被部署为一种受管学习服务,而非一组孤立的训练任务。在网络化边缘环境中,相互依赖的FL服务流必须在服务水平完成要求下协调异构客户端、非独立同分布(non-IID)数据、波动的通信延迟以及具有优先级约束的任务。这些耦合因素使得参与者管理对达到目标的时间和训练稳定性都至关重要。本文提出A-CoDa,一种用于管理相互依赖FL流的自适应聚类协调框架。A-CoDa首先使用基于标签分布散度(LDD)的贪心平衡聚类来构建统计上连贯且感知规模的客户端组,作为可扩展的管理抽象。在此基础上,我们设计了FedMIX,一种不确定性感知的簇内/簇间参与机制,通过损失-延迟-不确定性效用对客户端进行排序,并根据训练进度和延迟条件自适应控制跨簇探测。随后,一个依赖感知的DAG调度器协调层间任务执行,使得并行性和优先级约束同时得到满足。我们进一步提供了收敛性分析,将结果表述为充分的损失域设计界限,明确将可达到的误差下限和足够的通信轮数与LDD引起的采样不匹配、残余分布偏移、局部SGD漂移、随机方差以及自适应探测预算联系起来。在手写、可穿戴传感、产品图像和医学图像任务上的实验评估了A-CoDa在依赖FL工作流下的性能,并展示了其在减少端到端完成时间同时保持竞争性准确性方面的有效性。

英文摘要

Federated learning (FL) is increasingly deployed as a managed learning service rather than as a set of isolated training jobs. In networked edge environments, dependent FL service flows must coordinate heterogeneous clients, non-IID data, fluctuating communication latency, and precedence-constrained tasks under service-level completion requirements. These coupled factors make participant management central to both time-totarget performance and learning stability. This paper proposes A-CoDa, an adaptive clustered coordination framework for managing dependent FL flows. A-CoDa first uses label-distribution divergence (LDD)-based greedy-balanced clustering to construct statistically coherent and size-aware client groups, which serve as a scalable management abstraction. Building on this structure, we design FedMIX, an uncertainty-aware intra-/inter-cluster participation mechanism that ranks clients by a loss-latency-uncertainty utility and adaptively controls cross-cluster probing according to training progress and latency conditions. A dependency-aware DAG scheduler then orchestrates layer-wise task execution so that parallelism and precedence constraints are jointly respected. We further provide a convergence analysis that frames the result as a sufficient loss-domain design bound, explicitly relating the attainable error floor and sufficient communication rounds to LDDinduced sampling mismatch, residual distribution shift, local-SGD drift, stochastic variance, and adaptive probing budgets. Experiments on handwriting, wearable-sensing, product-image, and medical-imaging tasks evaluate A-CoDa under dependent FL workflows and demonstrate its effectiveness in reducing end-toend completion time while maintaining competitive accuracy.

Comments15 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑