发表机构
McMaster University; École de Technologie Supérieure(麦克马斯特大学; 高等技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对监督学习中噪声标签问题,提出XMix框架,利用自监督特征空间的局部平滑性,估计噪声率,识别干净样本并平衡类别选择,生成可靠伪标签,在极端噪声环境下性能优于现有方法。
AI 中文摘要
监督深度学习模型依赖大量准确标注的数据集,但噪声标注往往不可避免,在高噪声水平下会严重降低性能。近期的先进方法通过利用记忆效应的样本选择策略来解决此问题,但在极端噪声和类别不平衡情况下存在困难。我们提出XMix框架,利用自监督特征空间中的局部平滑性来系统地增强样本选择过程的各个阶段,不依赖可能有误的标签。通过估计噪声率、识别干净样本、确保类别平衡选择以及生成更可靠的伪标签,实验结果表明XMix在极端噪声环境中显著优于现有方法。
英文摘要
Supervised deep learning models rely on large, accurately labeled datasets, yet noisy annotations are often unavoidable and can severely degrade performance under high noise levels. Recent state-of-the-art methods tackle this by using sample selection strategies that exploit the memorization effect to filter out clean data for semi-supervised learning. However, these methods struggle with extreme noise, class imbalance, and require careful tuning or prior noise knowledge. To address these limitations, we propose XMix, a novel framework that leverages local smoothness in the self-supervised feature space to systematically enhance all stages of the sample selection process, without dependence on potentially corrupted labels. First, XMix estimates the noise rate using maximum likelihood among self-supervised feature neighbors. Second, these neighbors then help identify additional clean samples and ensure balanced selection across classes during sample selection. Finally, in the semi-supervised learning phase, XMix uses neighboring samples to generate more reliable pseudo-labels. Our empirical results show that XMix substantially outperforms existing methods in extremely noisy environments and maintains superior performance in standard LNL benchmarks.
DOI:10.1109/ICASSP55912.2026.11463541