arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MedFG-VQA:用于轻量级医学视觉问答的低频记忆与图注意力

MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

Haowen Gu, Gensheng Pei, Zeren Sun, Mingwu Ren, Xiangbo Shu, Yazhou Yao, Fumin Shen

arXiv 2608.26848首次发表:更新:

发表机构

School of Computer Science and Engineering, Nanjing University of Science and Technology; State Key Laboratory of Intelligent Manufacturing of Advanced Construction Machinery; Department of Electrical and Computer Engineering, Sungkyunkwan University; School of Computer Science and Engineering, University of Electronic Science and Technology of China(南京理工大学计算机科学与工程学院; 高端工程机械智能制造国家重点实验室; 成均馆大学电气与计算机工程系; 电子科技大学计算机科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MedFG-VQA是轻量级医学VQA框架,通过FMF和GACA实现高效视觉-文本对齐,构建含200万对问答的SynMed-VQA数据集,性能优于大模型且计算成本低,适用于临床部署。

AI 中文摘要

医学视觉问答(Med-VQA)对临床决策支持具有重要价值,但面临标注数据有限以及现有大型视觉语言模型计算需求高的挑战。我们提出MedFG-VQA,这是一种轻量级框架,利用记忆库增强基于DCT的低频特征,并采用图增强交叉注意力实现有效的视觉-文本对齐。具体而言,我们的方法包含两个关键组件:频率-记忆融合(FMF),通过从基于DCT分解构建的可学习记忆库中检索来增强低频特征;图感知交叉注意力(GACA),通过交叉注意力对齐视觉-文本特征并通过图卷积聚合对其进行优化。为解决数据稀缺问题,我们构建了SynMed-VQA,这是一个大规模合成数据集,包含超过200万对问答,涵盖9种成像模态和10个主要器官,由GPT-4o生成。在SynMed-VQA及其他三个标准生物医学VQA基准上的大量实验表明,MedFG-VQA与大得多的模型相比取得了相当或更优的性能,同时保持了显著更低的计算成本,凸显了其效率和临床部署的潜力。

英文摘要

Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language models. We propose MedFG-VQA, a lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment. Specifically, our approach features two key components: Frequency-Memory Fusion (FMF), which enhances low-frequency features by retrieving from a learnable memory bank built on DCT decomposition, and Graph-Aware Cross-Attention (GACA), which aligns visual-textual features via cross-attention and refines them through graph-convolutional aggregation. To address data scarcity, we construct SynMed-VQA, a large-scale synthetic dataset comprising over 2 million question-answer pairs across 9 imaging modalities and 10 major organs, generated with GPT-4o. Extensive experiments on SynMed-VQA and three other standard biomedical VQA benchmarks demonstrate that MedFG-VQA achieves competitive or superior performance compared to much larger models while maintaining significantly lower computational costs, highlighting its efficiency and potential for clinical deployment.

CommentsAccepted by CVPR 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑