arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RadSight:迈向感知可靠的多模态放射学图像理解

RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

Jianqin Liu, Weiwei Cao, Wanxing Chang, Ruifeng Yuan, Bowen Shi, Zhilin Zheng, Xianjie Zhang, Ling Zhang, Peng Wang, Jianpeng Zhang

arXiv 2607.22293首次发表:更新:

发表机构

DAMO Academy, Alibaba Group; Hupan Lab; University of Electronic Science and Technology of China(达摩院,阿里巴巴集团; 湖畔实验室; 电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对医学多模态大语言模型视觉解释可靠性低的问题,引入Perception-Bench基准分析问题。提出基于双2D/3D编码器架构的RadSight模型,经渐进课程学习训练,在多维度评估中优于现有模型,凸显低级视觉感知对可靠临床理解的关键作用。

AI 中文摘要

医学多模态大语言模型(MLLMs)越来越被期望执行复杂的图像理解任务,但其可靠性常因视觉解释中的频繁错误而受损。为系统追踪这些失败,我们从高级临床任务深入到基本视觉感知层次。为此引入Perception-Bench,一个含113万个样本的大规模基准,从六个维度评估医学MLLMs。分析发现现有MLLMs缺乏捕捉基本病变属性的能力。受此启发,提出RadSight,基于双2D/3D编码器架构,将医学图像理解设为四阶段渐进过程,用渐进课程学习在837万个面向感知的语料库上训练它。在Perception-Bench上,RadSight在所有六个评估维度上均优于现有MLLMs,在公共2D和3D医学基准上也有持续改进,证明强大的低级视觉感知是可靠临床理解的关键基础。代码和模型将公开可用。

英文摘要

Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often compromised by frequent errors in visual interpretation. To systematically trace these failures, we traverse the hierarchy from high-level clinical tasks down to fundamental visual perception. We therefore introduce Perception-Bench, a large-scale benchmark comprising 1.13 million samples that assesses medical MLLMs across six dimensions: attribute judgment, spatial grounding, spatial understanding, disease prediction, anomaly detection, and report generation, spanning both 2D and 3D radiology images. Our analysis on Perception-Bench reveals that existing MLLMs lack the ability to capture even the most basic lesion attributes, such as location, size, and density. This inability to ground clinical outputs in primary visual evidence reveals that the models' diagnostic unreliability is rooted in a critical but overlooked bottleneck in low-level visual perception. Motivated by this, we propose RadSight, a perception-driven MLLM built upon a dual 2D/3D encoder architecture that preserves native imaging spatial structures. RadSight formulates medical image understanding as a four-stage progressive process: visual-language alignment, fine-grained visual perception, clinical diagnosis, and diagnostic interpretation. The model is trained on an 8.37 million perception-oriented corpus using progressive curriculum learning. On Perception-Bench, RadSight consistently outperforms existing MLLMs across all six evaluation dimensions, with particularly strong gains in spatial grounding and clinical diagnosis. It also achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable clinical understanding. Code and model will be publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑