LiNC:基于逐样本置信度与高斯混合模型的轻量级噪声校正
LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling
浏览论文内容
中文总结 AI 辅助
针对医学成像数据集的标签噪声问题,本文提出LiNC方法,通过逐样本置信度参数结合高斯混合模型校正标签,在MedMNISTv2数据集上取得稳定的准确率提升且开销极小。
中文摘要 AI 辅助
标签噪声在医学成像数据集中十分常见,其产生原因包括评分者间差异、标注错误以及样本本身存在歧义等,这会严重破坏使用此类数据集训练的机器学习模型的可靠性与临床有效性。为应对这一挑战,本文提出了轻量级噪声校正(Lightweight Noise Correction, LiNC)方法,该方法在标准训练循环中为每个训练样本添加一个可训练的置信度参数,用于学习何时使用观测标签、何时让模型进行预测。其核心思路是使用观测标签与模型自身预测分布的凸组合进行训练,该组合由逐样本置信度参数控制。研究表明,该目标函数的梯度会在训练初期使干净样本与噪声样本的置信度值向相反方向变化,从而产生可分离的置信度分布。本文在置信度值上使用3分量高斯混合模型,将样本划分为干净、歧义、噪声三类,随后对噪声样本执行短期软校正阶段,最后进行最终硬校正阶段。在MedMNISTv2的10个2D数据集上,当标签噪声比例高达50%时,实验结果显示LiNC在准确率和错误标注检测能力上均取得了稳定提升;LiNC带来的渐近开销可忽略不计,其训练时间复杂度仍由基础网络主导,额外内存随训练集规模线性增长。
英文摘要
Label noise is common in medical imaging datasets due to factors such as inter-rater variability, annotation errors, and ambiguous cases. This can severely undermine the reliability and clinical effectiveness of machine learning models trained using those datasets. To address this challenge, we introduce Lightweight Noise Correction (LiNC), which adds a single trainable trust parameter per training sample and learns when to use the observed label and when to defer to the model during a standard training loop. The key idea is to train using a convex combination of the observed label and the model's own predictive distribution, controlled by a per-sample trust parameter. We show that the gradient of this objective drives trust values in opposite directions for clean versus noisy samples in the early training phase, yielding separable trust distributions. We use a 3-component Gaussian Mixture Model over the trust values to separate them into clean, ambiguous, and noisy cases and then execute a short soft-correction phase on the noisy cases and a final hard correction phase. Experiments on ten 2D datasets from MedMNISTv2 under label noise of up to 50% show consistent gains in accuracy and strong mislabel detection. LiNC adds negligible asymptotic overhead: the training-time complexity remains dominated by the base network, with additional memory growing linearly with the size of the training set.
发表机构
- University of Toronto(多伦多大学)
- The Hospital for Sick Children(病童医院)
- UHN KITE Research Institute(大学健康网络KITE研究所)
- Vector Institute(向量研究所)
- Institute of Biomedical Engineering(生物医学工程研究所)
- Rehabilitation Sciences Institute(康复科学研究所)
- Department of Laboratory Medicine and Pathobiology(检验医学与病理生物学系)
机构由 AI 辅助整理,请以论文原文为准。