arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MedUP:唤醒医学视觉-语言模型中的统一理解与感知能力

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

Yuan Wang, Hualiang Wang, Yixin Chen, Songtao Jiang, Shujian Gao, Jiaming Lin, Siming Fu, Jian Wu, Zuozhu Liu

arXiv 2608.10635首次发表:更新:

发表机构

Zhejiang University; Fudan University(浙江大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出MedUP模型,通过UniMedTok分词器在共享token空间统一医学视觉-语言模型的感知与理解,构建相关训练语料库和评估基准,在多医学任务上性能优于现有模型。

AI 中文摘要

医学视觉-语言模型(Med-VLMs)在将视觉内容转化为文本描述方面表现出色,但精准的视觉感知、分割和定位仍具挑战性。现有方法要么将区域表述为坐标字符串,要么依赖外部模块,这些模块将感知与理解解耦,造成了区域-语言对齐的表示差距。本文提出MedUP,一种在共享token空间中原生统一感知与理解的Med-VLM。其核心是UniMedTok,一种区域分词器,它将掩码编码为LLM词汇表中的离散token,使模型能够无缝地将掩码token与文本交织。我们构建了UniMed-Train,一个包含184万样本的语料库,涵盖文本引导分割、区域定位理解、医学视觉问答(VQA)和基于思维链(CoT)的分割任务,并引入UniMed-Bench用于统一评估。大量实验表明,MedUP在所有任务上均优于原生、智能体(agentic)和双解码器(dual-decoder)Med-VLMs,同时与专业分割器表现相当,证明了统一理解与感知建模的巨大潜力。

英文摘要

Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modules that decouple perception from understanding, creating representation gaps for region-language alignment. We present MedUP, a Med-VLM that natively unifies perception and understanding within a shared token space. At its core lies UniMedTok, a region tokenizer that encodes masks as discrete tokens in the LLM vocabulary, enabling the model to seamlessly interleave mask tokens with text. We curate UniMed-Train, a 1.84M-instance corpus spanning text-guided segmentation, region-grounded understanding, medical VQA and CoT-based segmentation, and introduce UniMed-Bench for unified evaluation. Extensive experiments show that MedUP outperforms native, agentic, and dual-decoder Med-VLMs across all tasks while remaining competitive with specialist segmentors, demonstrating the strong potential of unified understanding and perception modeling.

Comments10 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑