发表机构
DAMO Academy, Alibaba Group; College of Computer Science and Technology, Zhejiang University; Hupan Lab; School of Public Health, Zhejiang University(阿里巴巴达摩院; 浙江大学计算机科学与技术学院; 湖畔实验室; 浙江大学公共卫生学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对3D医学图像问答中多序列视觉冗余问题,提出SeVeR框架,引入BreMRIs-VQA基准,通过压缩体积为原型并检索互补证据,实现性能提升且减少视觉标记使用。
AI 中文摘要
体积医学视觉问答(VQA)需要对冗长且冗余的3D视觉标记序列进行推理,尤其是在多序列MRI中,互补模态提供了多样的诊断线索,但也使解码器暴露于大量重复的解剖区域。为研究多序列视觉冗余下的推理,我们首先引入BreMRIs-VQA,这是一个经临床整理的乳腺MRI基准,包含来自71.0K序列和12.9K患者的119万条问答对,涵盖自由文本和多项选择题。我们进一步提出SeVeR,一个选择性视觉暴露框架,该框架将密集体积压缩为模态级原型,并在解码期间通过感知变化的门控注意力检索互补的多级证据,采用边际效用自一致性目标进行训练,以抑制无用的检索。在BreMRIs-VQA和公共基准上的实验表明,SeVeR在显著减少视觉标记暴露量的同时,提升了判别式和生成式性能。
英文摘要
Volumetric medical VQA requires reasoning over long and redundant 3D visual token sequences, especially in multi-sequence MRI where complementary modalities provide diverse diagnostic cues but expose the decoder to many repeated anatomical regions. To investigate reasoning under multi-sequence visual redundancy, we first introduce BreMRIs-VQA, a clinically curated breast MRI benchmark with 1.19M QA pairs from 71.0K sequences and 12.9K patients, covering both free-text and multiple-choice questions. We further propose SeVeR, a selective visual exposure framework that compresses dense volumes into modality-wise prototypes and retrieves complementary multi-level evidence with change-aware gated attention during decoding, trained with a marginal-utility self-consistency objective that suppresses unhelpful retrieval. Experiments on BreMRIs-VQA and public benchmarks show that SeVeR improves both discriminative and generative performance while exposing substantially fewer visual tokens.