两阶段混合LoRA用于多任务医学视觉-语言学习
Two-Stage Mixture-of-LoRA for Multi-Task Medical Vision-Language Learning
- Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出基于MedGemma-1.5-4B的两阶段混合LoRA框架,通过共享与任务特定专家LoRA及两阶段训练,解决多任务医学视觉-语言学习中的任务冲突和数据不平衡,在FLARE 2026任务3上取得优异性能。
AI中文摘要:
医学视觉-语言模型(VLMs)允许单个模型执行从诊断分类到报告生成的临床图像分析任务。然而,联合适应受到异构输出格式、冲突的任务梯度和不平衡的训练数据的挑战。因此,我们提出了两阶段混合LoRA(Two-Stage Mixture-of-LoRA),这是一个基于MedGemma-1.5-4B构建的框架。该框架采用共享-特定混合LoRA架构,包含一个共享LoRA和六个任务特定的专家LoRA,并配合两阶段训练流程。在第一阶段,我们在所有任务上联合训练共享LoRA和所有任务特定的专家LoRA。在第二阶段,我们首先冻结骨干网络、共享LoRA和所有非目标专家,然后一次精调一个任务专家。分类和回归任务随后接受额外的模态平衡延续训练,其中较小的模态组被重复以匹配最大的组。在FLARE 2026任务3测试集上,所提出的方法在分类任务上达到0.85的平衡准确率,在多标签分类上达到0.48的微F1分数,检测F1为0.79,回归MAE为17.39。代码可在该HTTPS URL获取。
英文摘要:
Medical vision-language models (VLMs) allow a single model to perform clinical image analysis tasks ranging from diagnosis classification to report generation. However, joint adaptation is challenged by heterogeneous output formats, conflicting task gradients, and imbalanced training data. Hence, we present \textbf{Two-Stage Mixture-of-LoRA}, a framework built on MedGemma-1.5-4B. The framework uses a shared-specific Mixture-of-LoRA architecture comprising one shared LoRA and six task-specific expert LoRAs, together with a two-stage training procedure. In Stage 1, we jointly train the shared LoRA and all task-specific expert LoRAs on all tasks. In Stage 2, we first freeze the backbone, the shared LoRA, and all non-target experts, and refine one task expert at a time. Classification and regression then receive an additional modality-balanced continuation, in which smaller modality groups are repeated to match the largest group. In the FLARE 2026 Task 3 test sets, the proposed method achieves 0.85 balanced accuracy for classification, 0.48 micro-F1 for multi-label classification, 0.79 detection F1, and 17.39 regression MAE. Code is available at https://github.com/YuanYL03/MICCAI-FLARE-2026-Challenge-Task3-2D.