AI 中文总结
该研究以HAS-FL为案例,发现自适应子模型联邦学习中更新差异估计受容量混淆影响,自适应分配存在失效模式,且其收益源于容量预算与覆盖,而非分配智能。
AI 中文摘要
子模型联邦学习允许资源受限的客户端训练全局模型的宽度缩减版本,但现有方法仅根据设备资源分配容量。已有研究多次提出,下一步自然的方向是根据服务器已观测到的更新所估计的各客户端数据异质性来分配容量。我们以自适应容量分配框架HAS-FL为测试案例,探究这一步骤是否可行,研究结果分为三部分:其一,在可复现划分的真实标签分布差异验证下,对客户端异质性的更新差异估计受容量而非数据主导——在两个经校正的估计器、多个数据集及所有随机种子下,估计值与设备容量呈强负相关,控制容量后无数据信号残留,这一此前未被记录的混淆效应会影响任何从子模型更新中估计客户端统计量的方法;其二,自适应分配存在隐藏失效模式:当所有客户端的宽度均被限制在全宽度以下时,未被覆盖的参数保持随机初始化状态,会逐步破坏全局模型,简单的覆盖保证可消除该失效模式,也解释了为何均匀分配会失效;其三,匹配预算控制明确了自适应的作用:在相同平均预算下随机分配在两个图像基准上表现无差异,在自然划分的文本基准上,自适应策略是三种策略中最弱的,却消耗最多容量。子模型训练仍有价值,因为它能让受约束客户端以二次缩减的成本参与,但保护准确率的是参数覆盖而非分配智能,其表面收益来自容量预算与覆盖,未来设计需能分离容量效应的异质性信号。
英文摘要
Sub-model federated learning lets resource-constrained clients train width-reduced versions of a global model, but existing methods allocate capacity by device resources alone. A natural next step, allocating capacity by each client's data heterogeneity as estimated from the updates the server already observes, is suggested by recent methods that size sub-models from training-derived signals. We ask whether that step is possible, using HAS-FL, an adaptive capacity-allocation framework, as a test case. First, validated against ground-truth label-distribution divergence on reproducible partitions, update-divergence estimates of client heterogeneity are dominated by capacity rather than data: on both image benchmarks and every seed, the estimates correlate strongly and negatively with device capacity, and once capacity is controlled for their association with data heterogeneity is near zero or negative. Any method estimating client statistics from sub-model updates is exposed to this previously undocumented confound. Second, adaptive allocation has a hidden failure mode: when every client is capped below full width, the uncovered parameters stay at random initialization and progressively corrupt the global model. A simple coverage guarantee removes the failure and explains why uniform allocation collapses. Third, a matched-budget control settles what adaptivity contributes: random allocation to the same average budget matches the adaptive policy to within seed-to-seed variation on both image benchmarks, and on the naturally partitioned text benchmark the adaptive policy is the weakest of the three strategies while consuming the most capacity. Sub-model training admits constrained clients at quadratically reduced cost but gives up substantial accuracy relative to full-model training, and what protects that accuracy is parameter coverage rather than allocation intelligence.