arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PIcsC:分区诱导协变量偏移校正

PIcsC: Partitioning-Induced Covariate Shift Correction

Behraj Khan, Behroz Mirza, Syed Ahmad Chan Bukhari, Tahir Qasim Syed

arXiv 2607.25441首次发表:更新:

发表机构

Institute of Business Administration Karachi; Habib University; St. John’s University(卡拉奇工商管理学院; 哈比卜大学; 圣约翰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对跨训练数据分区的协变量偏移问题,提出PIcsC框架,利用Fisher信息矩阵近似分区散度并纳入正则化,还引入条件适应机制,实验表明该方法在多数据集及联邦学习基准上有效减轻协变量偏移,提升性能。

AI 中文摘要

跨训练数据分区的协变量偏移会在交叉验证、终身学习和联邦学习中使模型选择和参数估计产生偏差。我们提出了分区诱导协变量偏移校正(PIcsC),这是一个基于Fisher信息的正则化框架,可减轻数据分区与参考分布之间的分布不匹配。PIcsC使用Fisher信息矩阵(FIM)近似分区散度,并在优化过程中将结果统计量作为正则化项纳入。相同的公式适用于集中分区数据集(批次或交叉验证折)和固有分布式数据(联邦客户端或分散节点),仅需要分区局部梯度统计量而非原始数据。我们还引入了一种条件适应机制,将FIM偏移与KL散度相结合以检测显著的分布偏移,并仅在必要时激活正则化。在40多个数据集上的实验表明,在自然和合成协变量偏移下均有持续改进。在碎片化批次和折设置中,PIcsC分别将碎片化导致的性能下降降低了20%以上和25%以上。在七个联邦学习基准上,它始终比FedAvg、FedProx和SCAFFOLD高出3-5个百分点,且无需客户端特定的个性化。这些结果表明,Fisher信息为减轻集中式和分布式学习中分区诱导的协变量偏移提供了一种有效且统一的机制。

英文摘要

Covariate shift across training-data partitions biases model selection and parameter estimation in cross-validation, lifelong learning, and federated learning. We propose \textit{Partition-Induced Covariate-shift Correction} (\texttt{PIcsC}), a Fisher information-based regularization framework that mitigates distribution mismatch between data partitions and a reference distribution. \texttt{PIcsC} approximates partition divergence using the Fisher Information Matrix (FIM) and incorporates the resulting statistic as a regularizer during optimization. The same formulation applies to both centrally partitioned datasets (batches or cross-validation folds) and inherently distributed data (federated clients or decentralized nodes), requiring only partition-local gradient statistics rather than raw data. We further introduce a conditional adaptation mechanism that combines FIM shift with KL divergence to detect significant distribution shifts and activates regularization only when necessary. Experiments on more than 40 datasets demonstrate consistent improvements under both natural and synthetic covariate shift. On fragmented batch and fold settings, \texttt{PIcsC} reduces fragmentation-induced performance degradation by more than 20\% and 25\%, respectively. On seven federated learning benchmarks, it consistently outperforms FedAvg, FedProx, and SCAFFOLD by 3 -5 percentage points without requiring client-specific personalization. These results demonstrate that Fisher information provides an effective and unified mechanism for mitigating partition-induced covariate shift across both centralized and distributed learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑