面向资源高效联邦知识蒸馏的自适应异构压缩
Adaptive Heterogeneous Compression for Resource-Efficient Federated Knowledge Distillation
- School of Information Engineering, Guangdong University of Technology(广东工业大学信息工程学院)
- School of Communication and Information Engineering, Chongqing University of Posts and Telecommunications(重庆邮电大学通信与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对联邦知识蒸馏中梯度传输开销高、客户端资源与模型特性异质性未被充分利用的问题,提出ASCEND算法,通过多臂老虎机自适应选择压缩策略,在保持精度的同时降低通信与训练时间。
AI中文摘要:
联邦学习(FL)支持隐私保护的分布式模型训练,但面临异构模型架构和网络边缘有限通信资源的挑战。联邦知识蒸馏(FedKD)通过结合原型级参数聚合与异构模型间的知识迁移缓解模型异构问题,然而梯度传输仍会带来可观的通信开销,且现有压缩方法通常对客户端采用统一策略,忽略了其多样的模型特性与资源容量。为解决该问题,本文提出一种面向FedKD的异构压缩框架,使每个客户端可从候选策略集中选择压缩策略。我们将压缩策略选择问题建模为非平稳随机多臂老虎机(MAB)问题,其中每个臂对应一种压缩策略;设计一种效率感知奖励,同时考虑局部优化改进、全局知识对齐与执行时间。基于该建模,我们开发了自适应异构联邦知识蒸馏压缩算法(ASCEND),该算法采用指数移动平均(EMA)增强的ε-贪心策略以平衡探索与利用。在多个数据集上的实验结果表明,ASCEND可有效适配异构模型与资源设置,在保持竞争力模型精度的同时降低了通信开销与训练时间。
英文摘要:
Federated learning (FL) enables privacy-preserving distributed model training but faces challenges from heterogeneous model architectures and limited communication resources at the network edge. Federated knowledge distillation (FedKD) alleviates model heterogeneity by combining prototype-wise parameter aggregation and knowledge transfer across heterogeneous models. However, transmitting gradients still introduces considerable communication overhead, while existing compression approaches typically apply a uniform strategy across clients and ignore their diverse model characteristics and resource capacities. To address this issue, we propose a heterogeneous compression framework for FedKD that enables each client to select a compression strategy from a candidate strategy set. We formulate the compression strategy selection problem as a non-stationary stochastic multi-armed bandit (MAB), where each arm corresponds to a compression strategy. An efficiency-aware reward is designed by jointly considering local optimization improvement, global knowledge alignment, and execution time. Based on this formulation, we develop an Adaptive heterogeneouS Compression algorithm for fEderated kNowledge Distillation (ASCEND), which employs an exponential moving average (EMA)-enhanced $ε$-greedy policy to balance exploration and exploitation. Experimental results on multiple datasets demonstrate that ASCEND effectively adapts to heterogeneous model and resource settings, reducing communication overhead and training time while maintaining competitive model accuracy.