AI 中文总结
该研究提出统一医学像素-语言模型MedPixel,引入44万样本的MedPLG-440K数据集,通过联合多任务微调与像素级偏好优化训练,支持多类医学任务,性能优异且具备零样本迁移与鲁棒性。
AI 中文摘要
可靠的医学图像理解需要模型将临床语言和视觉推理与像素级定位关联起来。然而,医学视觉-语言模型往往缺乏精确的定位能力,而医学分割模型通常依赖明确的目标类别或精确的空间提示。这种差距因监督信号不匹配而加剧:分割数据集提供精确的掩码,但几乎没有语言监督;而医学视觉-语言数据很少将语言与密集空间注释配对。为解决这一差距,我们提出MedPixel,一种围绕共享语言-掩码接口构建的统一医学像素-语言模型。为提供可扩展的监督,我们引入MedPLG-440K,包含约44万个像素-语言任务样本,这些样本通过临床驱动的合成过程构建,无需外部大型语言模型(LLM)注释。MedPixel经过联合多任务监督微调训练,随后采用像素级偏好优化,该方法使用真实掩码作为离线验证器,从掩码质量中推导响应偏好。MedPixel支持广泛的任务,涵盖显式定位、隐式推理、空间交互、基于定位的解释以及医学视觉问答(VQA)。在这些任务中,MedPixel在像素级预测和响应生成方面均取得优异性能,同时具备对外部定位基准的有效零样本迁移能力,以及对不完美空间提示的鲁棒性。代码和模型检查点将在此httpsURL发布。
英文摘要
Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segmenters typically rely on explicit target categories or precise spatial prompts. This divide is reinforced by a supervision mismatch: segmentation datasets provide precise masks but little language supervision, whereas medical vision-language data rarely pair language with dense spatial annotations. To address this gap, we present MedPixel, a unified medical pixel-language model built around a shared language--mask interface. To provide scalable supervision, we introduce MedPLG-440K, comprising approximately 440K pixel-language task samples constructed through a clinically motivated synthesis process without external LLM annotation. MedPixel is trained with joint multi-task supervised fine-tuning followed by Pixel-Level Preference Optimization, which uses ground-truth masks as offline verifiers to derive response preferences from mask quality. MedPixel supports a broad spectrum of tasks spanning explicit grounding, implicit reasoning, spatial interaction, grounded explanation, and medical VQA. Across this task spectrum, MedPixel achieves strong performance in both pixel-level prediction and response generation, together with effective zero-shot transfer to external grounding benchmarks and robustness to imperfect spatial prompts. Code and model checkpoints will be released at https://github.com/yhy-whu/Medpixel.