AI 中文总结
针对医学图像序分类任务的标准混合增强会破坏序结构的问题,提出DisMix框架通过双码本VQ-VAE解耦序与非序特征,实现序感知混合,在四个医学数据集上的综合性能优于多个基线方法,且在数据稀缺等场景下仍有效。
AI 中文摘要
图像混合增强(image mixup)是一种广泛采用的数据增强策略,但它不适用于医学疾病分级这类序分类任务,这类任务的标签编码了严重程度的递进关系。标准混合增强会不加区分地将疾病严重程度线索(序特征)与外观层面的变化(非序特征)混合,生成的样本会破坏支撑临床严重程度分级的序结构。本文提出DisMix,一种用于序分类的序感知混合增强框架。DisMix通过双码本向量量化变分自编码器(dual-codebook VQ-VAE)解耦序特征与非序特征,允许各子空间独立混合:对序编码进行插值以生成有意义的中间等级,同时调整非序编码以引入外观多样性,且不会破坏序信号。在四个医学图像数据集上,DisMix在六个图像混合增强基线搭配六个序分类器的组合中表现出最佳的综合性能,且在数据稀缺和临床分级变异的情况下仍保持有效性。
英文摘要
Image mixup is a widely adopted data augmentation strategy, yet it is ill-suited for ordinal classification tasks such as medical disease grading, where labels encode a progression of severity. By indiscriminately blending disease-severity cues (ordinal) with appearance-level variation (non-ordinal), standard mixup produces samples that distort the very ordinal structure that underpins clinical severity grading. We introduce DisMix, an order-aware mixup framework for ordinal classification. DisMix disentangles ordinal and non-ordinal features via a dual-codebook VQ-VAE, allowing each subspace to be mixed independently: ordinal codes are interpolated to produce meaningful intermediate ranks, while non-ordinal codes are varied to introduce appearance diversity without corrupting the ordinal signal. Across four medical imaging datasets, DisMix shows the best aggregate performance among six image mixup baselines paired with six ordinal classifiers and remains effective under data scarcity and clinical grading variability.