发表机构
JDH Algo, JD Health International Inc.(京东健康国际有限公司 JDH Algo)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出参数解耦训练框架,通过分离<SEG>提示、两阶段指令微调及梯度缩放,实现医学推理与分割的统一,兼顾分割精度与推理性能。
AI 中文摘要
医学多模态大语言模型(MLLMs)日益期望不仅能回答临床问题,还能定位其预测背后的视觉证据。一种常见策略通过特殊的<SEG>令牌将视觉-语言模型(VLM)与SAM风格的分割连接起来,然而这种统一架构的全参数训练是困难的,因为图像级推理和像素级分割对共享表示空间提出了不同的要求。为解决这一问题,我们提出了一种用于统一医学推理与分割的参数解耦训练框架。该框架将<SEG>隐藏状态视为掩码解码器的语义到空间提示,并鼓励其与通用语言状态分离,从而减少模糊的分割提示和对推理表示的潜在干扰。它首先进行医学浅层对齐,在不干扰LLM的情况下使视觉特征适应临床语言;然后,受控指令微调塑造可分离的<SEG>提示状态,并通过Davies-Bouldin指数(DBI)进行监控,同时缩放进入语言主干的分割梯度;最后,在冻结VLM的情况下专门化SAM分支,以提高掩码精度而不改变推理参数。在医学指称分割、接地、视觉问答和文本问答基准上的实验表明,我们的框架实现了强大的语言条件分割,同时保持了有竞争力的推理能力。消融研究表明,两阶段指令微调、梯度缩放和分割专门化都对模型有所贡献。
英文摘要
Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision--language model (VLM) with SAM-style segmentation through a special <SEG> token, yet full-parameter training of this unified architecture is difficult because image-level reasoning and pixel-level segmentation impose different requirements on the shared representation space. To address this issue, we propose a parameter-decoupled training framework for unified medical reasoning and segmentation. The framework treats the <SEG> hidden state as a semantic-to-spatial prompt for the mask decoder and encourages it to become separable from generic language states, reducing ambiguous segmentation prompts and potential disruption to reasoning representations. It first performs medical shallow alignment to adapt visual features to clinical language without disturbing the LLM; then controlled instruction tuning shapes separable <SEG> prompt states, monitored by the Davies--Bouldin Index (DBI), while scaling segmentation gradients entering the language backbone; finally, the SAM branch is specialized with the VLM frozen to improve mask precision without altering reasoning parameters. Experiments on medical referring segmentation, grounding, visual QA, and textual QA benchmarks show that our framework achieves strong language-conditioned segmentation while preserving competitive reasoning ability. Ablations show that two-phase instruction tuning, gradient scaling, and segmentation specialization all contribute to the model.