发表机构
School of Computer Science and Engineering, Nanjing University of Science and Technology; State Key Laboratory of Intelligent Manufacturing of Advanced Construction Machinery; Department of Electrical and Computer Engineering, Sungkyunkwan University; School of Computer Science and Engineering, University of Electronic Science and Technology of China(南京理工大学计算机科学与工程学院; 高端工程机械智能制造国家重点实验室; 成均馆大学电气与计算机工程系; 电子科技大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MedFG-VQA是轻量级医学VQA框架,通过FMF和GACA实现高效视觉-文本对齐,构建含200万对问答的SynMed-VQA数据集,性能优于大模型且计算成本低,适用于临床部署。
AI 中文摘要
医学视觉问答(Med-VQA)对临床决策支持具有重要价值,但面临标注数据有限以及现有大型视觉语言模型计算需求高的挑战。我们提出MedFG-VQA,这是一种轻量级框架,利用记忆库增强基于DCT的低频特征,并采用图增强交叉注意力实现有效的视觉-文本对齐。具体而言,我们的方法包含两个关键组件:频率-记忆融合(FMF),通过从基于DCT分解构建的可学习记忆库中检索来增强低频特征;图感知交叉注意力(GACA),通过交叉注意力对齐视觉-文本特征并通过图卷积聚合对其进行优化。为解决数据稀缺问题,我们构建了SynMed-VQA,这是一个大规模合成数据集,包含超过200万对问答,涵盖9种成像模态和10个主要器官,由GPT-4o生成。在SynMed-VQA及其他三个标准生物医学VQA基准上的大量实验表明,MedFG-VQA与大得多的模型相比取得了相当或更优的性能,同时保持了显著更低的计算成本,凸显了其效率和临床部署的潜力。
英文摘要
Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language models. We propose MedFG-VQA, a lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment. Specifically, our approach features two key components: Frequency-Memory Fusion (FMF), which enhances low-frequency features by retrieving from a learnable memory bank built on DCT decomposition, and Graph-Aware Cross-Attention (GACA), which aligns visual-textual features via cross-attention and refines them through graph-convolutional aggregation. To address data scarcity, we construct SynMed-VQA, a large-scale synthetic dataset comprising over 2 million question-answer pairs across 9 imaging modalities and 10 major organs, generated with GPT-4o. Extensive experiments on SynMed-VQA and three other standard biomedical VQA benchmarks demonstrate that MedFG-VQA achieves competitive or superior performance compared to much larger models while maintaining significantly lower computational costs, highlighting its efficiency and potential for clinical deployment.
CommentsAccepted by CVPR 2026