arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CRAD:面向去中心化异构联邦学习的类别级可靠性感知知识蒸馏

CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning

Baraa Bilbeisi, Mengchen Fan, Baocheng Geng, Qing Tian

arXiv 2609.00446首次发表:更新:

发表机构

University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对传统联邦学习的同质性假设局限,提出CRAD框架,通过类别级可靠性加权融合对等教师预测,在三类图像分类基准的异构非IID场景中实现更优全局准确率。

AI 中文摘要

传统联邦学习(FL)依赖参数平均,这要求客户端具备双重同质性:即相同的模型架构,且在非独立同分布(non-IID)数据下性能会下降。实际部署场景通常会打破这两项假设。我们通过构建去中心化知识蒸馏框架规避了这两个问题,在该框架中,每个客户端会基于自身本地数据评估对等方的模型快照,并从生成的软预测中进行知识蒸馏。由于知识通过共享的类别后验进行传递,客户端可自由运行不同的架构;且由于每个教师模型都在客户端自身设备上针对学生模型进行评估,原始数据不会离开客户端,无需中央服务器或公共数据集。在该设定下,我们识别并解决了一个未被充分研究的问题:如何融合对等教师的预测结果。现有方法(如均匀平均)忽略了教师与类别间知识可靠性的差异。我们提出类别级可靠性感知知识蒸馏(CRAD),其按类别先剔除与对等共识不一致的教师,再对剩余教师进行加权平均,权重为每个教师的类别级可靠性(精确率或方差的倒数)。由于n个样本的准确率方差随1/n缩放,支持度会自动纳入考量:在过滤后留存的教师中,教师对某一类别的信任度取决于其在该类别上的准确性与证据充分程度。在三个图像分类基准(CIFAR-10、CIFAR-100及PathMNIST结肠病理数据集)上,针对严重非IID偏置下的异构架构,CRAD在全局准确率上始终优于对比方法。

英文摘要

Conventional federated learning (FL) relies on parameter averaging, which forces clients to be doubly homogeneous: it demands an identical architecture and degrades under non-IID data. Real-world deployments usually break both assumptions. We sidestep both by building a decentralized knowledge distillation framework in which each client evaluates its peers' model snapshots on its own local data and distills from the resulting soft predictions. Because knowledge is transferred through the shared class posterior, clients are free to run different architectures; and because every teacher is evaluated on the student's own device, raw data never leaves the client, with no central server or public dataset required. Within this setting, we identify and address an under-examined problem: how to combine the peer teacher predictions. Existing methods, like uniform averaging, ignore how knowledge reliability varies across teachers and classes. We propose Class-wise Reliability-Aware Distillation (CRAD), which, per class, first discards teachers that disagree with the peer consensus and then takes a weighted average of the rest, weighting each teacher by its per-class reliability (precision, or inverse variance). Since the variance of an accuracy from $n$ samples scales as $1/n$, support enters automatically: among the teachers that survive filtering, a teacher is trusted for a class to the degree that it is both accurate and well-evidenced for it. On three image-classification benchmarks (CIFAR-10, CIFAR-100, and PathMNIST colon pathology), across heterogeneous architectures under severe non-IID skew, CRAD consistently outperforms competing methods in global accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑