QwenVLConnector:面向细粒度临床感知与文本生成的快速统一医学视觉语言模型聊天机器人
QwenVLConnector: A Fast, Unified Medical VLM Chatbot for Fine-Grained Clinical Perception and Text Generation
- University of Wisconsin - Madison(威斯康星大学麦迪逊分校)
- University of North Carolina - Chapel Hill(北卡罗来纳大学教堂山分校)
- Carnegie Mellon University(卡内基梅隆大学)
- Northwestern University(西北大学)
- Mayo Clinic, College of Medicine and Science(梅奥诊所医学与科学学院)
- University of Houston(休斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
QwenVLConnector基于Qwen2.5-VL,通过轻量密集多层连接器统一分类、检测、计数等结构化感知与报告生成,在FLARE-2D上显著提升检测F1至0.85,并支持无需参数更新的多模态上下文学习。
AI中文摘要:
大多数医学视觉语言模型(VLM)擅长开放式报告生成和视觉问答(VQA),但在统一界面内对结构化、细粒度的临床感知支持有限。我们提出了QwenVLConnector,一个基于Qwen2.5-VL的医学聊天机器人,它将分类、多标签分类、文本化检测、计数、回归和自由形式报告生成统一在单一的下一个词元目标下。我们的关键组件是一个轻量级密集多层连接器(Connector),它聚合低层和高层视觉特征,通过预训练的视觉合并器(Merger)对齐它们,并将其与最终视觉表示融合,而不增加序列长度。这种设计丰富了视觉词元的空间和语义线索,同时保持了效率。在FLARE-2D上,QwenVLConnector将检测F1从0.55提升到0.85,将单标签分类从0.37提升到0.51,并将报告生成的GREEN评分比Qwen2.5-VL基线提高了最多18.3分。我们进一步探索了用于报告生成的多模态上下文学习,显示出在不更新模型参数的情况下也能获得额外改进。总体而言,QwenVLConnector为结合结构化医学感知与开放式临床文本生成提供了一个统一且高效的框架。我们的代码可在https://这个URL找到。
英文摘要:
Most medical vision-language models (VLMs) excel at open-ended report generation and VQA but provide limited support for structured, fine-grained clinical perception within a unified interface. We present QwenVLConnector, a Qwen2.5-VL-based medical chatbot that unifies classification, multi-label classification, textualized detection, counting, regression, and free-form report generation under a single next-token objective. Our key component is a lightweight dense multi-layer Connector that aggregates low- and high-level visual features, aligns them through the pretrained vision Merger, and fuses them with the final visual representation without increasing sequence length. This design enriches visual tokens with complementary spatial and semantic cues while preserving efficiency. On FLARE-2D, QwenVLConnector improves detection F1 from 0.55 to 0.85, raises single-label classification from 0.37 to 0.51, and boosts report-generation GREEN by up to 18.3 points over the Qwen2.5-VL baseline. We further explore multimodal in-context learning for report generation, showing additional improvements without updating model parameters. Overall, QwenVLConnector offers a unified and efficient framework for combining structured medical perception with open-ended clinical text generation. Our code can be found at https://github.com/plnguyen2908/QwenConnector.