AI 中文总结
该研究针对VLMs视觉预测不可靠的频谱响应刚性问题,提出HAFI-VLM,通过任务条件频率通路与相关组件优化,在多个VLMs基准上提升了视觉感知性能。
AI 中文摘要
视觉语言模型(VLMs)在预测需要细粒度视觉证据时仍不可靠。我们发现一个此前被忽视的原因:频谱响应刚性。尽管不同图像和任务间存在显著的频率变化,但预训练视觉编码器表现出持续的、编码器特有的分层频谱轮廓,在下游微调下仅发生微小变化。由于预训练视觉编码器仅接收图像,它们无法将频谱提取适配当前查询所需的证据。因此,我们提出HAFI-VLM,它在保留预训练语义表示的同时引入了任务条件频率通路。分层自适应频率注入(HAFI)利用文本调制、空间对齐的交叉注意力,在多个编码器深度检索互补的低、中、高频证据;视觉增强层适配器进一步重新校准大型语言模型(LLM)的浅层注意力,以有效利用增强的视觉标记。在LLaVA-1.5和Qwen2.5-VL上的实验表明,该方法在通用视觉问答(VQA)、文本丰富理解和幻觉鲁棒性方面取得了一致提升,优于表示级增强方法和大多数基于分辨率或裁剪的方法,且无需额外高分辨率编码。机制分析显示,HAFI在保留语义注意力的同时恢复了任务相关的频谱分配,确立了频率增强是改善VLM感知的独特且有效途径。
英文摘要
Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spectral response rigidity. Despite substantial frequency variation across images and tasks, pretrained vision encoders exhibit persistent, encoder-specific layerwise spectral profiles that change only marginally under downstream fine-tuning. Since pretrained vision encoders only receive images, they cannot adapt spectral extraction to the evidence required by the current query. We therefore propose HAFI-VLM, which introduces a task-conditioned frequency pathway while preserving the pretrained semantic representation. Hierarchical Adaptive Frequency Injection (HAFI) retrieves complementary low-, mid-, and high-frequency evidence at multiple encoder depths using text-modulated, spatially aligned cross-attention. A Visual Enrichment Layer Adapter further recalibrates shallow LLM attention to effectively utilize the enriched visual tokens. Experiments on LLaVA-1.5 and Qwen2.5-VL demonstrate consistent improvements in general VQA, text-rich understanding, and hallucination robustness, outperforming representation-level enhancement methods and most resolution- or cropping-based approaches without additional high-resolution encoding. Mechanistic analyses show that HAFI restores task-dependent spectral allocation while retaining semantic attention, establishing frequency enrichment as a distinct and effective route for improving VLM perception.
Comments11 pages, 8 figure