MMDS:一种融合图像分析与基于知识的科室咨询的多模态医学诊断系统
MMDS: A Multimodal Medical Diagnosis System Integrating Image Analysis and Knowledge-based Departmental Consultation
- Xidian University(西安电子科技大学)
- Xijing Hospital, The Fourth Military Medical University(第四军医大学西京医院)
- Shaanxi University of Science and Technology(陕西科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出多模态医学诊断系统MMDS,通过专用多模态模型分析医学图像、面部情绪和面瘫视频,并结合科室知识库路由管理机制优化大语言模型的检索增强生成,从而提供更准确的专业诊断。
AI中文摘要:
我们提出MMDS,一个能够识别医学图像和患者面部细节,并提供专业医学诊断的系统。该系统由两个核心部分组成:第一部分是医学图像和视频分析。我们训练了一个专门的多模态医学模型,能够解读医学图像,并准确分析患者的面部情绪和面瘫状况。该模型在FER2013面部情绪识别数据集上达到了72.59%的准确率,其中对“快乐”情绪的识别准确率为91.1%。在面瘫识别方面,该模型达到了92%的准确率,比GPT-4o高出30%。基于该模型,我们开发了一个用于分析面瘫患者面部运动视频的解析器,实现了对面瘫严重程度的精确分级。在对30个面瘫患者视频的测试中,该系统展现出83.3%的分级准确率。第二部分是专业医学响应的生成。我们采用了一个大语言模型,并与医学知识库集成,基于对医学图像或视频的分析生成专业诊断。核心创新在于我们开发了一种科室特定知识库路由管理机制,在该机制中,大语言模型按医学科室对数据进行分类,并在检索过程中确定要查询的适当知识库。这显著提高了RAG(检索增强生成)过程中的检索准确性。
英文摘要:
We present MMDS, a system capable of recognizing medical images and patient facial details, and providing professional medical diagnoses. The system consists of two core components:The first component is the analysis of medical images and videos. We trained a specialized multimodal medical model capable of interpreting medical images and accurately analyzing patients' facial emotions and facial paralysis conditions. The model achieved an accuracy of 72.59% on the FER2013 facial emotion recognition dataset, with a 91.1% accuracy in recognizing the "happy" emotion. In facial paralysis recognition, the model reached an accuracy of 92%, which is 30% higher than that of GPT-4o. Based on this model, we developed a parser for analyzing facial movement videos of patients with facial paralysis, achieving precise grading of the paralysis severity. In tests on 30 videos of facial paralysis patients, the system demonstrated a grading accuracy of 83.3%.The second component is the generation of professional medical responses. We employed a large language model, integrated with a medical knowledge base, to generate professional diagnoses based on the analysis of medical images or videos. The core innovation lies in our development of a department-specific knowledge base routing management mechanism, in which the large language model categorizes data by medical departments and, during the retrieval process, determines the appropriate knowledge base to query. This significantly improves retrieval accuracy in the RAG (retrieval-augmented generation) process.