arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13939cs.CVcs.AI

CMCNet:将超声图像嵌入与文本TI-RADS表示对齐以用于细粒度甲状腺分类

CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li

AI总结:

本研究构建含600个甲状腺结节的STN数据集,提出CMCNet模型,通过中心间隔对比损失对齐图像与文本嵌入,实现细粒度甲状腺分类,性能优于多类基线方法。

AI中文摘要:

超声是评估甲状腺结节的主要成像模态,ACR TI-RADS框架通过五类超声特征标准化诊断,这些特征被汇总为五个风险等级(TR1-TR5)。尽管该框架在临床实践中被广泛采用,但大多数深度学习方法聚焦于良恶性二分类,而多分类预测及对特征级监督的明确利用仍未得到充分探索,这在很大程度上归因于标注数据有限。本研究构建了STN数据集,包含600个甲状腺结节的配对横纵超声图像、边界框标注以及全部五类TI-RADS特征的完整标签。遵循临床决策流程,研究探索结构化特征信息如何在训练时引导表示学习,而推理时仅需图像输入。研究发现,基于标准化特征描述生成的文本嵌入可作为TI-RADS风险等级的稳定替代表示。基于此观察,提出CMCNet,其通过中心间隔对比损失将图像嵌入与固定文本嵌入对齐,该损失同时促进类内紧凑性与类间分离性。实验结果表明,该嵌入对齐策略比直接多任务学习更具数据效率和鲁棒性,且在不平衡场景中,始终优于InfoNCE、中心损失、强多任务基线及VQA式多模态模型。该数据集可通过doi: https://doi.org/10.5281/zenodo.19125693免费获取,源代码可通过提供的链接获取。

英文摘要:

Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5). Although widely adopted in clinical practice, most deep learning approaches focus on binary malignancy classification, while multi-class prediction and explicit utilization of feature-level supervision remain underexplored, largely due to limited annotated data. In this study, we introduce the STN dataset of 600 thyroid nodules with paired transverse and longitudinal ultrasound images, bounding box annotations, and complete labels for all five TI-RADS feature categories. Following the clinical decision process, we investigate how structured feature information can guide representation learning during training while requiring only images at inference. We demonstrate that text embeddings derived from standardized feature descriptions form a stable surrogate representation for TI-RADS risk levels. Based on this observation, we propose CMCNet, which aligns image embeddings to fixed textual embeddings via a Center-Margin Contrastive Loss that simultaneously promotes intra-class compactness and inter-class separation. Experimental results show that this embedding alignment strategy is more data-efficient and robust than direct multitask learning, and consistently outperforms InfoNCE, center loss, a strong multitask baseline, and a VQA-style multimodal model, particularly in imbalanced settings. The dataset is freely available at doi: 10.5281/zenodo.19125693 and the source code is available at: https://www.healthinformaticslab.org/supp/.

↑