发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对少样本医学分类中标记数据稀缺,现有方法忽视样本对分类有用性衡量与优化的问题,提出类对比影响(C2I)准则,通过强化学习用C2I奖励微调扩散模型,引导生成类信息丰富样本,提高了下游准确性和鲁棒性。
AI 中文摘要
当标记数据稀缺时,现成的扩散模型可扩充训练集用于少样本医学图像分类,但并非所有生成样本对下游任务都同样有用。现有方法主要通过提高真实性、多样性或域适应性来改进合成数据,却忽略了一个更基本的问题:如何衡量和优化样本对分类的有用性?我们用类对比影响(C2I)解决此问题,它通过样本对分类器基于梯度的影响来量化样本有用性。我们发现有效样本呈现出强烈的C2I差距:其损失梯度与同一类别的验证梯度对齐,与其他类别的相反。我们的分析还表明,此类高C2I样本是有助于细化决策边界并提高鲁棒性的硬的、边界近端示例。基于此,我们使用基于C2I的奖励通过强化学习微调扩散模型,引导生成朝着类信息丰富的样本。在几个少样本医学成像基准上,C2I引导的生成比基于扩散的增强基线提高了下游准确性和鲁棒性,表明当由任务有用性而非仅图像质量引导时,合成增强最有效。
英文摘要
When labeled data are scarce, off-the-shelf diffusion models can augment training sets for few-shot medical image classification, but not all generated samples are equally useful for the downstream task. Existing approaches largely improve synthetic data by increasing realism, diversity, or domain adaptation, while overlooking a more fundamental question: how should sample usefulness for classification be measured and optimized? We address this with Class-Contrastive Influence (C2I), a criterion that quantifies a sample's usefulness through its gradient-based influence on the classifier. We find that effective samples exhibit a strong C2I gap: their loss gradients align with validation gradients from the same class and oppose those from other classes. Our analysis further suggests that such high-C2I samples are hard, boundary-proximal examples that help refine the decision boundary and improve robustness. Building on this insight, we fine-tune diffusion models with reinforcement learning using a C2I-based reward to steer generation toward class-informative samples. Across several few-shot medical imaging benchmarks, C2I-guided generation improves downstream accuracy and robustness over diffusion-based augmentation baselines, showing that synthetic augmentation is most effective when guided by task usefulness rather than image quality alone.