arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.09230cs.CVcs.AIcs.CLcs.LG

使用多模态大语言模型和消费级摄像头诊断肩部疾病

Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras

  • Bytedance(字节跳动)
  • Peking University(北京大学)
  • Peking University People’s Hospital(北京大学人民医院)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

Jindong Hong, Wencheng Zhang, Shiqin Qiao, Jianhai Chen, Jianing Qiu, Chuanyang Zheng, Qian Xu, Yun Ji, Qianyue Wen, Weiwei Sun, Hao Li, Huizhen Li, Huichao Wan… 展开作者

Jindong Hong, Wencheng Zhang, Shiqin Qiao, Jianhai Chen, Jianing Qiu, Chuanyang Zheng, Qian Xu, Yun Ji, Qianyue Wen, Weiwei Sun, Hao Li, Huizhen Li, Huichao Wang, Kai Wu, Meng Li, Yijun He, Lingjie Luo, Jiankai Sun

更新

AI总结:

本研究提出HMVDx框架,利用消费级摄像头视频和两个多模态大语言模型分别进行动作理解与疾病诊断,通过可用性指数评估,使肩关节损伤诊断准确率较直接视频诊断提升79.6%。

AI中文摘要:

肩部疾病,如冻结肩(又称粘连性关节囊炎),是影响全球人群健康的常见疾病,在老年人和从事重复性肩部工作的劳动者中发病率较高。在医疗资源匮乏的地区,实现早期准确诊断面临重大挑战,迫切需要低成本且易于扩展的辅助诊断方案。本研究引入由消费级设备拍摄的视频作为诊断基础,以降低用户成本。我们聚焦于多模态大语言模型(MLLMs)在肩部疾病初步诊断中的创新应用,并提出了一种混合运动视频诊断框架(HMVDx)。该框架将动作理解和疾病诊断两个任务分开,分别由两个MLLMs完成。除了传统的评估指标外,本工作还通过医学决策的逻辑过程(动作识别、运动诊断和最终诊断)提出了一种名为可用性指数的新指标。该指标从整个医学诊断路径的角度评估MLLMs在医学领域的有效性,揭示了低成本MLLMs在医疗应用中对医疗从业者的潜在价值。在实验对比中,与直接视频诊断相比,HMVDx在诊断肩关节损伤方面的准确率提高了79.6%,这对未来MLLMs在医学领域视频理解应用的研究是一项重要的技术贡献。

英文摘要:

Shoulder disorders, such as frozen shoulder (a.k.a., adhesive capsulitis), are common conditions affecting the health of people worldwide, and have a high incidence rate among the elderly and workers engaged in repetitive shoulder tasks. In regions with scarce medical resources, achieving early and accurate diagnosis poses significant challenges, and there is an urgent need for low-cost and easily scalable auxiliary diagnostic solutions. This research introduces videos captured by consumer-grade devices as the basis for diagnosis, reducing the cost for users. We focus on the innovative application of Multimodal Large Language Models (MLLMs) in the preliminary diagnosis of shoulder disorders and propose a Hybrid Motion Video Diagnosis framework (HMVDx). This framework divides the two tasks of action understanding and disease diagnosis, which are respectively completed by two MLLMs. In addition to traditional evaluation indicators, this work proposes a novel metric called Usability Index by the logical process of medical decision-making (action recognition, movement diagnosis, and final diagnosis). This index evaluates the effectiveness of MLLMs in the medical field from the perspective of the entire medical diagnostic pathway, revealing the potential value of low-cost MLLMs in medical applications for medical practitioners. In experimental comparisons, the accuracy of HMVDx in diagnosing shoulder joint injuries has increased by 79.6\% compared with direct video diagnosis, a significant technical contribution to future research on the application of MLLMs for video understanding in the medical field.

↑