arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

C-Score:开放世界未标记污染下半监督学习鲁棒性评估的超越准确率指标

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

Tsao-Lun Chen, Chi-Cheng Fu, Han-Yi E. Chou, Shun-Feng Su

arXiv 2608.20667首次发表:更新:

AI 中文总结

本文针对开放世界未标记污染下半监督学习的鲁棒性评估问题,提出C-Score框架,通过多维度指标揭示干净准确率无法检测的隐藏退化,验证了内部诊断信号的必要性。

AI 中文摘要

基于伪标签的半监督学习(SSL)因简洁性和可扩展性取得了优异性能,但通常基于封闭世界假设,即未标记数据与标记数据来自同一分布。在实际部署中,未标记数据常来自开放环境,可能包含分布外(OOD)样本,这些样本可能获得高置信度预测并被当作有效目标样本纳入训练,导致评估问题:即使干净分布内测试准确率看似稳定,SSL的内部学习动态可能已恶化。为解决此问题,本文从诊断评估视角研究开放世界未标记污染下基于伪标签的SSL中的隐藏崩溃,提出C-Score这一紧凑框架,在预测、特征表示、优化三个互补空间评估训练行为,包含用于未标记预测行为的PLE和CCI、用于偏离标记语义锚点的Sem-Drift、用于标记与未标记优化兼容性的Grad-Align。在CIFAR-10和CIFAR-100上,结合多种OOD源、不同污染比例及四种基于伪标签的SSL算法的实验显示,C-Score指标可揭示干净准确率无法检测的隐藏退化:在SVHN污染下,CCI上升超280%,而最佳准确率仍与无污染基线相差在3%以内;近OOD源(CIFAR-100、STL-10)导致最高14.9%的准确率崩溃(FlexMatch,r=0.5)。结果表明,仅干净准确率不足以评估开放世界中SSL的鲁棒性,内部诊断信号对未标记污染下更可靠的鲁棒性评估是必要的。

英文摘要

Pseudo-label-based semi-supervised learning has achieved strong performance due to its simplicity and scalability. However, it is typically developed under a closed-world assumption that unlabeled data are drawn from the same distribution as labeled data. In practical deployment, unlabeled data are often collected from open environments and may contain OOD samples. Under such contamination, OOD samples may still receive high-confidence predictions and be incorporated into training as if they were valid target examples. This creates an important evaluation problem: clean in-distribution test accuracy may appear stable even when the internal learning dynamics of SSL have already deteriorated. To address this issue, we study hidden collapse in pseudo-label-based SSL under open-world unlabeled contamination from a diagnostic evaluation perspective. We present C-Score, a compact framework that evaluates training behavior in three complementary spaces: prediction, feature representation, and optimization. C-Score includes PLE and CCI for unlabeled prediction behavior, Sem-Drift for deviation from labeled semantic anchors, and Grad-Align for the compatibility between labeled and unlabeled optimization. Experiments on CIFAR-10 and CIFAR-100 with multiple OOD sources, varying contamination ratios, and four pseudo-label-based SSL algorithms show that C-Score metrics reveal hidden degradation that clean accuracy alone fails to detect: under SVHN contamination, CCI rises over 280% while best-accuracy remains within 3% of the uncontaminated baseline; near-OOD sources (CIFAR-100, STL-10) cause up to 14.9% accuracy collapse (FlexMatch, r=0.5). The results suggest that clean accuracy alone is insufficient for evaluating SSL robustness in open-world environments, and that internal diagnostic signals are necessary for more reliable robustness assessment under unlabeled contamination.

CommentsAccepted at IEEE ICSSE 2026 for oral presentation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑