arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2501.16282eess.IVcs.AIcs.CV

Brain-Adapter:利用适配器微调多模态大语言模型增强神经系统疾病分析

Brain-Adapter: Enhancing Neurological Disorder Analysis with Adapter-Tuning Multimodal Large Language Models

  • The University of Texas at Arlington(德克萨斯大学阿灵顿分校)
  • Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校)
  • The University of Georgia(佐治亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Jing Zhang, Xiaowei Yu, Yanjun Lyu, Lu Zhang, Tong Chen, Chao Cao, Yan Zhuang, Minheng Chen, Tianming Liu, Dajiang Zhu

更新

AI总结:

本文提出Brain-Adapter方法,通过引入轻量级瓶颈层和CLIP策略对齐多模态数据,在低计算成本下显著提升了3D医学图像的神经系统疾病诊断准确率。

AI中文摘要:

理解大脑疾病对于准确的临床诊断和治疗至关重要。多模态大语言模型(MLLMs)的最新进展提供了一种在文本描述支持下解释医学图像的有前景的方法。然而,先前的研究主要集中在2D医学图像上,未能充分探索3D图像中更丰富的空间信息,且基于单一模态的方法因忽略其他模态中包含的关键临床信息而受到限制。为了解决这一问题,本文提出了Brain-Adapter,这是一种新颖的方法,它引入了一个额外的瓶颈层来学习新知识,并将其注入到原始预训练知识中。主要思路是引入轻量级瓶颈层以在训练更少参数的同时捕获关键信息,并利用对比语言-图像预训练(CLIP)策略在统一表示空间内对齐多模态数据。大量实验证明了我们方法在整合多模态数据方面的有效性,能够在不产生高计算成本的情况下显著提高诊断准确率,突出了增强现实世界诊断工作流的潜力。

英文摘要:

Understanding brain disorders is crucial for accurate clinical diagnosis and treatment. Recent advances in Multimodal Large Language Models (MLLMs) offer a promising approach to interpreting medical images with the support of text descriptions. However, previous research has primarily focused on 2D medical images, leaving richer spatial information of 3D images under-explored, and single-modality-based methods are limited by overlooking the critical clinical information contained in other modalities. To address this issue, this paper proposes Brain-Adapter, a novel approach that incorporates an extra bottleneck layer to learn new knowledge and instill it into the original pre-trained knowledge. The major idea is to incorporate a lightweight bottleneck layer to train fewer parameters while capturing essential information and utilize a Contrastive Language-Image Pre-training (CLIP) strategy to align multimodal data within a unified representation space. Extensive experiments demonstrated the effectiveness of our approach in integrating multimodal data to significantly improve the diagnosis accuracy without high computational costs, highlighting the potential to enhance real-world diagnostic workflows.

补充信息

↑