发表机构
Seoul National University (SNU); Institute of New Media and Communications, SNU(首尔大学; 首尔大学新媒体与通信研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无线分割学习中的通信瓶颈,提出重要性感知类别平衡稀疏化(ICS),利用Grad-CAM分数在服务器端筛选特征通道,降低客户端开销并缓解非独立同分布数据下的精度下降,实验验证其优于基线。
AI 中文摘要
无线分割学习(SL)通过将上层网络卸载到服务器来减少设备端计算量,然而每次迭代中传输高维中间特征仍然是主要的通信瓶颈。现有方法在客户端使用基于幅度、统计量或聚类等任务无关的标准来选择特征,这增加了客户端的处理负担,并且在非独立同分布(non-i.i.d.)的客户端数据下往往导致精度下降。我们提出了一种重要性感知的类别平衡稀疏化(ICS)方法,这是一种轻量级方法,服务器在反向传播过程中利用基于Grad-CAM的、从真实类别logit获得的分数对特征通道进行排序。各类别的分数被聚合成一个类别平衡的、与标签无关的重要性向量,以缓解标签分布偏移下的头部类别偏差,每个客户端在下一轮复用该向量以保留前N个特征通道,且不会增加额外的客户端前向或反向传播。我们进一步推导了一个非渐近收敛界,该界隔离了稀疏化引入的误差,并刻画了在固定通信预算下稀疏化比例和小批量大小如何共同影响收敛性,同时我们分析了ICS相对于代表性基线的通信和计算开销。除了基于CNN的串行分割学习,我们将ICS扩展到并行分割学习和基于Transformer的模型。实验表明,ICS始终优于基线方法,在严重的非独立同分布数据划分下提升更为显著。
英文摘要
Wireless split learning (SL) reduces on-device computation by offloading upper layers to a server, yet transmitting high-dimensional intermediate features at each iteration remains a major communication bottleneck. Existing methods select features at the client side using task-agnostic criteria such as magnitude, statistics, or clustering, which increases client-side processing and often degrades accuracy under non-independent and identically distributed (non-i.i.d.) client data. We propose importance-aware class-balanced sparsification (ICS), a lightweight approach in which the server ranks feature channels using Grad-CAM-based scores obtained from the true-class logit during backpropagation. The per-class scores are aggregated into a class-balanced, label-agnostic importance vector that mitigates head-class bias under label skew, and each client reuses this vector in the next round to retain the top-$N$ feature channels, incurring no additional client-side forward or backward passes. We further derive a non-asymptotic convergence bound that isolates the sparsification-induced error and characterizes how the sparsification ratio and mini-batch size jointly affect convergence under a fixed communication budget, and we analyze the communication and computational overhead of ICS against representative baselines. Beyond sequential CNN-based SL, we extend ICS to parallel split learning and to transformer-based models. Experiments show that ICS consistently outperforms the baselines, with larger gains under severe non-i.i.d. partitions.
CommentsAccepted for publication in IEEE Journal on Selected Areas in Communications