arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结合影像与临床数据的可解释多模态深度学习用于口腔潜在恶性病变检测

Explainable Multimodal Deep Learning Integrating Imaging and Clinical Data for Oral Potentially Malignant Disorder Detection

Ruilin You, Yihan Wang, Jiabin Chen, Cherie Wink, Petra Wilder-Smith, Rongguang Liang, Bofan Song

arXiv 2609.04512首次发表:更新:

发表机构

Wyant College of Optical Sciences, University of Arizona; Beckman Laser Institute & Medical Clinic, University of California Irvine(亚利桑那大学怀恩特光学科学学院; 加州大学尔湾分校贝克曼激光研究所与医学诊所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发了多模态深度学习框架M2-OPMDNet,整合口腔影像与临床数据,可准确检测口腔潜在恶性病变,优于单模态方法,且具备可解释性,为口腔癌筛查提供支持。

AI 中文摘要

口腔潜在恶性病变(OPMDs)是口腔癌的关键前驱病变,但因表型异质性大且与良性病变重叠,临床检测颇具挑战性。尽管基于影像的深度学习在自动筛查方面展现潜力,但在依赖患者特异性风险因素的真实场景中,仅靠视觉信息可能不足。本研究开发了M2-OPMDNet,这一多模态深度学习框架将配准后的白光口腔内影像与自体荧光口腔内影像,结合结构化临床信息用于OPMD检测。研究设计了定制问卷,以标准化、可重复的格式采集临床相关风险因素与症状,用于整合影像衍生特征。采用前瞻性收集的、反映真实世界筛查场景的数据集,评估了多种影像编码器,包括传统卷积神经网络和基于基础模型的架构。利用SHapley加性解释(SHAP)评估模型可解释性,以量化特征级和模态级贡献。M2-OPMDNet的AUC达到0.952,优于单模态方法,且对视觉上细微的病变性能更优。SHAP分析显示,结构化临床变量对风险估计贡献显著,可补充影像特征。这些结果表明,结合白光与自体荧光影像及结构化临床数据的可解释多模态学习,可提供准确、透明且符合临床实际的OPMD检测,M2-OPMDNet为真实世界口腔癌筛查与决策支持提供了可扩展框架。

英文摘要

Oral potentially malignant disorders (OPMDs) are critical precursors to oral cancer, yet clinical detection remains challenging because of substantial phenotypic heterogeneity and overlap with benign conditions. Although image-based deep learning shows promise for automated screening, visual information alone may be insufficient in real-world settings, where diagnostic decisions also rely on patient-specific risk factors. We developed M2-OPMDNet, a multimodal deep learning framework that integrates co-registered white-light and autofluorescence intraoral images with structured clinical information for OPMD detection. A customized questionnaire was designed to capture clinically relevant risk factors and symptoms in a standardized, reproducible format for integration with image-derived features. Multiple image encoders, including conventional convolutional neural networks and foundation model-based architectures, were evaluated using a prospectively collected dataset reflecting real-world screening conditions. Model interpretability was assessed using SHapley Additive exPlanations (SHAP) to quantify feature- and modality-level contributions. M2-OPMDNet achieved an AUC of 0.952, outperforming unimodal approaches and showing improved performance for visually subtle lesions. SHAP analysis demonstrated that structured clinical variables contributed substantially to risk estimation and complemented imaging features. These results demonstrate that explainable multimodal learning combining white-light and autofluorescence imaging with structured clinical data can provide accurate, transparent, and clinically grounded OPMD detection. M2-OPMDNet offers a scalable framework for real-world oral cancer screening and decision support.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑