面向多模态拆分学习的贡献感知带宽分配
Contribution-Aware Bandwidth Allocation for Multimodal Split Learning
- Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对多模态拆分学习的上行链路带宽分配问题,提出ModalShare分配器,利用Shapley贡献分数设置各模态保留率,在5倍压缩下显著提升了CREMA-D和MVSA数据集的准确率,且适配多种压缩器、数据集和预算。
AI中文摘要:
多模态模型正日益成为网络边缘感知的默认选项,但它们几乎全部在数据中心训练,因为持有多个传感器流的客户端无法为每个模态托管一个编码器。拆分学习通过仅在设备上保留第一层使此类训练可行,代价是上行链路必须在每一步传输每个模态的压缩激活值。现有压缩方案为每个模态分配相同的保留率,因此共享预算按压缩激活维度分配,而该量与每个模态对融合预测的贡献无关。我们将这种分配明确为决策,称之为模态间分配:在固定上行链路预算下,每种策略传输的预期有效载荷相同,仅在该有效载荷跨模态的分配方式上存在差异。我们的分配器ModalShare根据服务器在已接收激活值联盟上计算的Shapley贡献分数设置每个模态的保留率。测量该分数不会增加上行链路流量,也不会增加客户端侧计算,且无需预先知道哪个流对应哪个模态。在5倍压缩的匹配有效载荷下,ModalShare在CREMA-D和MVSA数据集上的准确率比相同保留率分别提高15.4和12.4个百分点,在三种压缩器、三个数据集和四个预算下均表现出强劲性能。我们发现现有压缩器在多模态场景中表现不佳,而ModalShare可恢复其遗留的部分增益。
英文摘要:
Multimodal models are increasingly the default option for perception at the network edge, yet they are trained almost entirely in the datacenter, because a client holding several sensor streams cannot host an encoder per modality. Split Learning makes such training feasible by keeping only the first layers on the device, at the cost of an uplink that must carry smashed activations for every modality at every step. Existing compression schemes give each modality the same keep-ratio, so the shared budget is divided in proportion to smashed-activation dimension, a quantity unrelated to how much each modality contributes to the fused prediction. We make that division an explicit decision and call it inter-modality allocation: under a fixed uplink budget, every policy transmits the same expected payload and differs only in how that payload is split across modalities. Our allocator, ModalShare, sets each modality's keep-ratio from a Shapley contribution score that the server computes over coalitions of activations it has already received. Measuring this score adds no uplink traffic and no client-side computation, and needs no prior knowledge of which stream is which. ModalShare improves accuracy over equal keep-ratios by 15.4 and 12.4 percentage points on CREMA-D and MVSA at matched payload in 5x compression, with strong performance across three compressors, three datasets, and four budgets. We show that existing compressors underperform in multimodal settings, with ModalShare recovering what gains are left behind.