发表机构
College of Traditional Chinese Medicine, Hebei University(河北大学中医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过对中医舌诊数据集开展消融实验,明确了深度学习模型的关键设计原则,为相关多标签医学图像分类任务提供了可推广的依据。
AI 中文摘要
深度学习在中医(TCM)舌诊自动化方面已展现潜力,但相关设计空间仍未得到充分探索。我们在TongueDx2数据集(含5109张图像,其中976张经专家标注)及11101个样本的合并数据集上,通过严格的5折交叉验证开展了涵盖20余个模型版本的系统消融研究,对比了6种骨干网络架构、4种损失函数、5种数据增强策略和6种训练策略。基于976个样本的最优模型采用ConvNeXt-Tiny结合受限数据增强与弱组集成,加权F1值达0.6625;基于11101个样本的最优模型加权F1值达0.7761。研究得出6项关键设计原则:(1)ConvNeXt-Tiny具备最优参数效率;(2)二值交叉熵(BCE)较非对称损失(Asymmetric Loss)性能提升2.7%;(3)受限颜色增强至关重要;(4)弱组集成替换较概率平均性能提升2.1%;(5)数据规模扩展带来20.6%的性能提升;(6)标签维度从13扩展至45会引发灾难性性能崩溃(加权F1值从0.78降至0.22)。这些原则可推广至类别不平衡的多标签医学图像分类任务。
英文摘要
Deep learning has shown promise for automated tongue diagnosis in traditional Chinese medicine (TCM), yet the design space remains underexplored. We conducted a systematic ablation study spanning 20+ model versions under rigorous 5-fold cross-validation on TongueDx2 (5,109 images, 976 expert-annotated) and a merged dataset of 11,101 samples. We compared six backbone architectures, four loss functions, five augmentation strategies, and six training strategies. The best 976-sample model achieved weighted-F1 of 0.6625 using ConvNeXt-Tiny with restrained augmentation and weak-group ensemble, while the best 11,101-sample model reached weighted-F1 of 0.7761. Six key design principles emerged: (1) ConvNeXt-Tiny offers optimal parameter efficiency; (2) BCE substantially outperforms Asymmetric Loss (+2.7%); (3) restrained color augmentation is critical; (4) weak-group ensemble replacement (+2.1%) outperforms probability averaging; (5) data scaling yielded +20.6% improvement; (6) expanding from 13 to 45 label dimensions caused catastrophic collapse (0.78 to 0.22). These principles are generalizable to multi-label medical image classification with class imbalance.
Comments30 pages, 8 figures, 9 tables