arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于头颈癌PET/CT解读的专用大型多模态模型

A Specialized Large Multimodal Model for Interpreting PET/CT in Head and Neck Cancer

Haengbok Chung, SunGyu Kim, Joo hyun Lee, Sangjin Bae, Min Jeong Cho, Minseok Suh, Jae Sung Lee

arXiv 2609.05532首次发表:更新:

发表机构

Seoul National University; Seoul National University College of Medicine; Seoul National University Hospital; Seoul National University Graduate School; Seoul National University Bundang Hospital; Brightonix Imaging Inc.(首尔大学; 首尔大学医学院; 首尔大学医院; 首尔大学研究生院; 首尔大学盆唐医院; Brightonix Imaging 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过两级课程微调LLaVA-NeXT构建专用LMM,用于头颈癌PET/CT自动解读,在外部验证中显著优于通用模型,支持临床诊断和医学教育。

AI 中文摘要

背景:使用PET/CT诊断头颈癌在临床上具有挑战性且耗时,因为该区域的解剖结构复杂,这促使了计算机辅助诊断(CAD)的发展。通用大型多模态模型(LMMs)在医学背景下仍受限于领域特定知识不足、隐私和安全问题以及冗长性,这促使了专用独立LMM的发展。目的:我们使用大规模多机构PET/CT数据集、定制训练课程和自回归训练,评估了专用LMM在头颈癌自动化PET/CT解读中的可行性。方法:使用两级课程对LLaVA-NeXT进行微调,图像-对话对由两位放射科医生从公共数据中策划。数据集包含临床重要标注,如原发肿瘤存在和转移性淋巴结位置。第1级使用28,000个图像-对话对学习基本信息,包括模态类型和高代谢。第2级使用12,975个图像-对话对学习原发肿瘤存在以及颈部淋巴结转移的存在和解剖位置。外部验证包括四个具有不同成像设备的机构。结果:专用LMM显著优于ChatGPT和LLaVA-NeXT。在第2级外部验证中,ROUGE-L、ROUGE-S、余弦相似度、精确率、召回率和F1分别为0.8751、0.8794、0.8324、0.8794、0.8711和0.8751,而通用模型始终得分低于0.1。原发肿瘤分类准确率内部为83.14±1.15%,外部为69.03±0.81%。对于淋巴结定位,相应得分分别为0.6389、0.6257、0.5287、0.5782、0.6371和0.6648。结论:专用LMM在快速、准确的基于PET/CT的诊断支持和医学教育方面显示出有前景的结果,突显了其临床转化的潜力。

英文摘要

Background: Diagnosing head and neck cancer using PET/CT is clinically challenging and time-consuming due to the anatomical complexity of the region, motivating computer-aided diagnosis (CAD). Generalist Large Multimodal Models (LMMs) remain limited in medical contexts by insufficient domain-specific knowledge, privacy and security concerns, and verbosity, motivating specialized standalone LMMs. Purpose: We evaluated the feasibility of a specialized LMM for automated PET/CT interpretation in head and neck cancer using a large-scale multi-institutional PET/CT dataset, a tailored training curriculum, and autoregressive training. Methods: LLaVA-NeXT was fine-tuned using a two-level curriculum with image-conversation pairs curated by two radiologists from public data. The dataset included clinically important annotations such as primary tumor presence and metastatic lymph node location. Level 1 used 28,000 image-conversation pairs to learn basic information, including modality type and hypermetabolism. Level 2 used 12,975 pairs to learn primary tumor presence and the existence and anatomical location of cervical lymph node metastases. External validation included four institutions with diverse imaging devices. Results: The specialized LMM substantially outperformed ChatGPT and LLaVA-NeXT. In Level-2 external validation, ROUGE-L, ROUGE-S, Cosine Similarity, Precision, Recall, and F1 were 0.8751, 0.8794, 0.8324, 0.8794, 0.8711, and 0.8751, while generalist models consistently scored below 0.1. Primary tumor classification accuracy was 83.14 +/- 1.15% internally and 69.03 +/- 0.81% externally. For lymph node localization, the corresponding scores were 0.6389, 0.6257, 0.5287, 0.5782, 0.6371, and 0.6648. Conclusion: Specialized LMMs show promising results for fast, accurate PET/CT-based diagnostic support and medical education, highlighting their potential for clinical translation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑