arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NeoRed:面向新生儿呼吸系统疾病诊断的知识-逻辑对齐多模态大语言模型

NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis

Yinan Liu, Hongtai Xia, Haoran Xu, Jiankang Hong, Yu Jianli, Jingkuan Song, Ye Luo

arXiv 2609.03527首次发表:更新:

发表机构

School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对现有多模态大语言模型在新生儿呼吸系统疾病诊断中存在的领域差距与上下文整合不足问题,构建了首个定制化MLLM NeoRed,通过KLA框架优化诊断,在新生儿数据集上表现优于现有模型,且在成人基准上性能具竞争力。

AI 中文摘要

新生儿呼吸系统疾病是导致新生儿发病和死亡的主要原因,给临床实践带来重大挑战。尽管近年来取得了诸多进展,现有多模态大语言模型(MLLM)在新生儿诊断领域面临两大关键局限:(1)训练数据以成人数据为主导致的领域差距;(2)多维度临床上下文整合不足,难以实现准确诊断。为应对这些挑战,我们收集了两个真实临床数据集NeoCXR和NeoCXR-EV,并提出了NeoRed——据我们所知,首个专为新生儿呼吸系统疾病定制的MLLM,填补了新生儿诊断报告生成领域的空白。为增强对异质性临床上下文与胸部X光片的联合诊断能力,我们设计了一种新型知识-逻辑对齐(KLA)框架,该框架从三个维度约束模型行为:1)知识先验注入(KPI)将新生儿科医生的诊断先验融入多模态表征,引导跨模态的疾病特异性注意力;2)诊断逻辑约束(DLC)使生成报告的语义与多模态诊断逻辑对齐;3)视觉语义对齐(VSA)建立视觉特征与影像结论之间的语义对应关系。大量实验表明,NeoRed可实现准确的新生儿诊断报告生成,在NeoCXR数据集上达到53.29%的ROUGE-L值和65.19%的临床效能F1值,优于现有MLLM。NeoRed在成人基准数据集MIMIC-CXR和IU-Xray上也保持了具有竞争力的报告生成性能。数据集将在申请后提供。

英文摘要

Neonatal respiratory diseases are a major cause of neonatal morbidity and mortality, posing substantial challenges in clinical practice. Despite recent advances, existing Multimodal Large Language Models (MLLMs) face two key limitations in neonatal diagnosis: (1) domain gap arising from predominantly adult training data; (2) insufficient integration of multidimensional clinical context for accurate diagnosis. To address these challenges, we collect two real-world clinical datasets (NeoCXR and NeoCXR-EV) and propose NeoRed, to the best of our knowledge, the first MLLM tailored for neonatal respiratory disease, filling the gap in neonatal diagnostic reports generation. To enhance joint diagnosis from heterogeneous clinical context and chest X-rays, we design a novel Knowledge-Logic-Alignment (KLA) framework which constrains model behavior from three perspectives: 1) Knowledge Prior Injection (KPI) incorporates neonatologist-inspired diagnostic priors into multimodal representations, guiding disease-specific attention across modalities; 2) Diagnostic Logic Constraint (DLC) aligns the semantics of generated reports with multimodal diagnostic logic; and 3) Visual Semantic Alignment (VSA) establishes semantic correspondence between visual features and imaging conclusions. Extensive experiments demonstrate that NeoRed enables accurate neonatal diagnostic reports generation, achieving ROUGE-L of 53.29% and Clinical Efficacy F1 score of 65.19% on NeoCXR, outperforming existing MLLMs. NeoRed also preserves competitive report generation performance on adult benchmarks (MIMIC-CXR and IU-Xray). Datasets will be available upon application.

Comments9 pages 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑