AI 中文总结
提出频域双分支融合模块,结合BiomedCLIP与BioBART,在PMC-VQA预训练后于VQA-RAD、SLAKE微调,提升医学VQA性能且架构轻量高效。
AI 中文摘要
医学视觉问答(VQA)需要将细微的视觉证据,包括病灶纹理、边界清晰度及弥散密度变化,与临床语言对齐。现有在空间域操作的多模态融合方法可能无法充分利用视觉和文本表示中存在的互补频率信息。我们提出一种双分支频域融合模块,该模块以输入问题为条件进行频谱滤波,可自适应选择全局低频结构和细粒度高频细节,再重建空间表示以生成答案。为提供更丰富的滤波频谱,我们从冻结的BiomedCLIP编码器的早期纹理敏感层和最终语义层提取互补特征,并在使用BioBART解码器进行分阶段联合训练前,用对称InfoNCE目标将这两种特征与问题表示对齐。我们在PMC-VQA上预训练所提模型,并在VQA-RAD和SLAKE基准上进行微调,结果表明,频域感知的多模态融合可提升医学VQA性能,同时保持轻量高效的架构。
英文摘要
Medical Visual Question Answering (VQA) requires aligning subtle visual evidence, including lesion texture, boundary sharpness, and diffuse density changes, with clinical language. Existing multimodal fusion approaches operating in the spatial domain may not fully exploit complementary frequency information present in visual and textual representations. We introduce a dual-branch frequency-domain fusion module that conditions spectral filtering on the input question, enabling adaptive selection of global low-frequency structure and fine-grained high-frequency detail before reconstructing the spatial representation for answer generation. To provide a richer spectrum for filtering, we extract complementary features from early texture-sensitive and final semantic layers of a frozen BiomedCLIP encoder and align both with the question representation using a symmetric InfoNCE objective prior to staged joint training with a BioBART decoder. We pretrain the proposed model on PMC-VQA and fine-tune it on the VQA-RAD and SLAKE benchmarks, demonstrating that frequency-aware multimodal fusion improves medical VQA performance while maintaining a lightweight and efficient architecture.
Comments7 Pages, 4 figures, under review at AAAI 27