arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31454cs.LGcs.AIcs.CV

不同的损坏,不同的信号:联邦数据质量中的不确定性与损失

Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality

  • University of Glasgow(格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

Bradley Scott, Zeqi Luo, Edmond S. L. Ho

AI总结:

本文在联邦学习中比较了输入条件不确定性与预测-标签损失两种损坏检测信号,发现标签翻转宜用损失检测,图像噪声宜用期望熵不确定性,强调信号需与损坏类型匹配。

AI中文摘要:

联邦学习(FL)中的数据损坏可能影响输入或标签,但目前尚不清楚输入条件不确定性和预测-标签损失是否能同样有效地暴露这些损坏模式。本文比较了FL中两种损坏检测信号:输入条件不确定性和预测-标签损失。不确定性信号通过学习的随机(aleatoric)方差估计以及蒙特卡洛(MC)dropout方差和熵度量来表征,而损失则根据提供的标签计算。我们针对加性图像噪声和持续随机标签翻转测试了这些信号。在ResNet-20上,使用CIFAR-10和SVHN数据集,在非独立同分布(non-IID)的Dirichlet分区下,两种损坏类型表现出不同的行为。对于持续随机标签翻转,客户端内每样本的接收者操作特征曲线下面积(AUC)在CIFAR-10上为0.85,在SVHN上为0.95(针对预测-标签损失),而所有不确定性估计器均停留在随机水平(0.49-0.50)。这一模式与模型尽管提供的标签错误但仍对底层图像保持自信的情况一致。对于图像噪声,期望熵不确定性高于随机水平(CIFAR-10上为0.67,SVHN上为0.66),而损失响应相当(两者均为0.64)。因此,每种信号对于不同的损坏类型是更强的检测器:预测-标签损失用于持续标签翻转,期望熵不确定性用于图像噪声,且其优势随着联邦范围内损坏流行率的增加而变得明显。稳健的联邦数据质量评估应将信号与损坏类型匹配,而不是仅依赖不确定性来应对所有损坏类型。

英文摘要:

Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL: input-conditional uncertainty and prediction-label loss. The uncertainty signal is characterised using a learned aleatoric variance estimate together with Monte Carlo (MC) dropout variance and entropy measures, while the loss is computed against the supplied label. We test these signals against additive image noise and persistent random label flips. On ResNet-20 with CIFAR-10 and SVHN under Dirichlet partitions with data that are not independent and identically distributed (non-IID), the two corruption types behave differently. For persistent random label flips, the within-client per-sample area under the receiver operating characteristic curve (AUC) is 0.85 on CIFAR-10 and 0.95 on SVHN for prediction-label loss, while every uncertainty estimator stays at chance (0.49--0.50). This pattern is consistent with the model remaining confident in the underlying image despite the supplied label being wrong. For image noise, expected-entropy uncertainty rises above chance (0.67 on CIFAR-10 and 0.66 on SVHN), while loss responds comparably (0.64 on both). Each signal is therefore the stronger detector for a different corruption: the prediction-label loss for persistent label flips, and expected-entropy uncertainty for image noise, with its advantage becoming apparent as federation-wide corruption prevalence increases. Robust FL data-quality assessment should match the signal to the corruption rather than rely on uncertainty alone across corruption types.

补充信息

↑