发表机构
Florida International University; Security, Optimization, and Learning for InterDependentnetworks laboratory (solid lab); The Pennsylvania State University; Florida Institute for Human and Machine Cognition (IHMC)(佛罗里达国际大学; 安全、优化与相互依赖网络学习实验室(solid lab); 宾夕法尼亚州立大学; 佛罗里达人类与机器认知研究所(IHMC))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对客户端内外部数据混合的复合异质性,提出FedSEE方法,通过任务对齐路由将输入分配给专用专家,利用凸规划恢复任务专家,实验显示性能提升2.9点,最差服务四分位数提升3.7点。
AI 中文摘要
在协作式基础模型微调中,客户端数据很少是同质的。相反,客户端通常拥有不同数据分布或任务的未知混合。传统联邦学习主要解决客户端之间的异质性,而没有显式解决每个客户端内部的潜在任务混合。我们将此设置研究为复合异质性,即数据在客户端之间和客户端内部都是异质的。我们研究在共享的冻结表示上的自适应,并表明,当任务共享相同的特征几何时,在平方损失下,客户端任务混合的最优模型是其底层任务最优模型的凸组合。因此,单个本地训练的模型代表客户端的整体任务混合,而单个输入可能来自不同的底层任务分布。这促使将输入路由到专用专家,并且我们表明,当任务最优形成一个单纯形时,对于真正混合的客户端,任务对齐的路由比任何单个自适应模型实现更低的风险。借助一小部分带任务标签的公共样本,我们推导出一个凸规划来恢复任务专家并将其与对应任务匹配。我们的路由分析表明,有效的专业化需要与每个客户端的任务混合对齐的输入相关专家选择。受此分析启发,我们提出了FedSEE。在我们的实验中,FedSEE避免了评估基线中观察到的负迁移,并在总体上提高了2.9个点,在最差服务四分位数上提高了3.7个点。
英文摘要
In collaborative foundation model fine-tuning, client data is rarely homogeneous. Instead, clients typically possess unknown mixtures of distinct data distributions, or tasks. Conventional federated learning primarily addresses heterogeneity across clients without explicitly resolving latent task mixtures within each client. We study this setting as compound heterogeneity, where data is heterogeneous both across and within clients. We study adaptation over a common frozen representation and show that, when tasks share the same feature geometry, the optimal model for a client's task mixture under squared loss is a convex combination of the optimal models for its underlying tasks. Thus, a single locally trained model represents the client's overall task mixture, while individual inputs may be drawn from different underlying task distributions. This motivates routing inputs to specialized experts, and we show that, when the task optima form a simplex, task-aligned routing achieves lower risk than any single adapted model for genuinely mixed clients. With access to a small set of task-labeled public samples, we derive a convex program to recover task experts and match them to their corresponding tasks. Our routing analysis shows that effective specialization requires input-dependent expert selection aligned with each client's task mixture. Motivated by this analysis, we propose FedSEE. Across our experiments, FedSEE avoids the negative transfer observed in the evaluated baselines and improves performance by 2.9 points overall and 3.7 points for the worst-served quartile.
Comments63 pages, 9 figures