arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SeVeR:面向3D医学图像问答的选择性视觉暴露与检索

SeVeR: Selective Visual Exposure and Retrieval for 3D Medical Image Question Answering

Yaojun Hu, Danyang Tu, Yang Liu, Jiajin Zhang, Wei Fang, Zhiqiang Liu, Chunlai Dong, Yingda Xia, Haochao Ying, Jian Wu, Ling Zhang

arXiv 2608.25630首次发表:更新:

发表机构

DAMO Academy, Alibaba Group; College of Computer Science and Technology, Zhejiang University; Hupan Lab; School of Public Health, Zhejiang University(阿里巴巴达摩院; 浙江大学计算机科学与技术学院; 湖畔实验室; 浙江大学公共卫生学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对3D医学图像问答中多序列视觉冗余问题,提出SeVeR框架,引入BreMRIs-VQA基准,通过压缩体积为原型并检索互补证据,实现性能提升且减少视觉标记使用。

AI 中文摘要

体积医学视觉问答(VQA)需要对冗长且冗余的3D视觉标记序列进行推理,尤其是在多序列MRI中,互补模态提供了多样的诊断线索,但也使解码器暴露于大量重复的解剖区域。为研究多序列视觉冗余下的推理,我们首先引入BreMRIs-VQA,这是一个经临床整理的乳腺MRI基准,包含来自71.0K序列和12.9K患者的119万条问答对,涵盖自由文本和多项选择题。我们进一步提出SeVeR,一个选择性视觉暴露框架,该框架将密集体积压缩为模态级原型,并在解码期间通过感知变化的门控注意力检索互补的多级证据,采用边际效用自一致性目标进行训练,以抑制无用的检索。在BreMRIs-VQA和公共基准上的实验表明,SeVeR在显著减少视觉标记暴露量的同时,提升了判别式和生成式性能。

英文摘要

Volumetric medical VQA requires reasoning over long and redundant 3D visual token sequences, especially in multi-sequence MRI where complementary modalities provide diverse diagnostic cues but expose the decoder to many repeated anatomical regions. To investigate reasoning under multi-sequence visual redundancy, we first introduce BreMRIs-VQA, a clinically curated breast MRI benchmark with 1.19M QA pairs from 71.0K sequences and 12.9K patients, covering both free-text and multiple-choice questions. We further propose SeVeR, a selective visual exposure framework that compresses dense volumes into modality-wise prototypes and retrieves complementary multi-level evidence with change-aware gated attention during decoding, trained with a marginal-utility self-consistency objective that suppresses unhelpful retrieval. Experiments on BreMRIs-VQA and public benchmarks show that SeVeR improves both discriminative and generative performance while exposing substantially fewer visual tokens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑