发表机构
China University of Petroleum (East China); Shanghai Jiao Tong University; Wuhan University; Great Wall Motor; Nanyang Technological University; The University of Hong Kong(中国石油大学(华东); 上海交通大学; 武汉大学; 长城汽车; 南洋理工大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MLLMs推理时的视觉表示退化问题,提出SSVAL方法,通过VAPI及辅助对齐损失实现稳定视觉锚,性能优于现有方法。
AI 中文摘要
尽管多模态大语言模型(MLLMs)已取得一定进展,但其在视觉感知方面仍存在缺陷。在视觉指令微调后,MLLM的内部表示在推理过程中会迅速偏离原始语义状态,导致严重的信息退化。现有方法尝试利用外部视觉基础模型(VFMs)来对齐内部表示,但我们发现直接与VFMs对齐虽能增强视觉语义,却无法缓解表示偏差。为解决该问题,我们提出空间-光谱视觉锚学习(SSVAL)。SSVAL的核心是视觉锚提示注入(VAPI),它在训练过程中引入能从外部VFMs吸收丰富知识的提示,使其在推理时可作为稳定的视觉锚以缓解表示偏差。此外,我们还加入辅助的空间和频域表示对齐损失,在LLM的中间层提供互补的视觉特定监督。大量实验表明,SSVAL的性能显著优于现有方法,代码可在项目页面获取。
英文摘要
Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instruction tuning, internal MLLM representations rapidly deviate from their original semantic states during inference, causing severe information degradation. While existing methods attempt to leverage external vision foundation models (VFMs) to align internal representations, we find that direct alignment with VFMs enhances visual semantics but fails to mitigate representation deviation. To address this, we propose Spatial-Spectral Visual Anchor Learning (SSVAL). The core of SSVAL is Visual Anchor Prompt Injection (VAPI), which introduces prompts that absorb rich knowledge from external VFMs during training, enabling them to serve as stable visual anchors that mitigate representation deviation during inference. Additionally, we incorporate auxiliary spatial and frequency-domain representation alignment losses to provide complementary vision-specific supervision at intermediate LLM layers. Extensive experiments demonstrate that SSVAL significantly outperforms existing methods. Code are available on our \href{https://msls38.github.io/SSVAL/}{project page}.
CommentsThis paper has been accepted by ACM MM 2026