AI 中文总结
本研究提出LENS方法,将VLM激活分解为局部低秩高斯邻域,揭示LLaVA与Qwen3-VL的不同模态融合轨迹,该方法可实现因果生成干预并提升多模态检索性能。
AI 中文摘要
视觉语言模型(VLMs)在共享残差流中处理图像块和文本 token,但两种模态交互的局部几何仍知之甚少。大多数可解释性方法识别全局线性方向,可能遗漏全局高维但局部低维的表示。我们提出LENS(Local Explanation of Neighborhood Subspaces,邻域子空间的局部解释),一种使用因子分析混合模型将VLM激活分解为局部低秩高斯邻域的方法。应用于LLaVA-1.5-7B和Qwen3-VL-8B,LENS揭示了与各模型融合机制一致的不同深度依赖融合轨迹:LLaVA在后续层逐步融合模态,而Qwen3-VL早期融合、部分重新分离并在输出附近重新组合。自动化多模态标注管道为这些邻域分配简洁语义描述。向邻域质心插值激活可在模态内和跨模态因果重定向生成,在多数评估条件下优于均值差和VL-SAE;在一项LLaVA视觉到视觉设置中,MFA达到VL-SAE分数的5.7倍。人工评估发现MFA干预与提示相当,显著强于其他干预基线。最后,MFA系数空间将Qwen3-VL在最深评估层的图像到渲染文本检索R@1从14.9%提升至48.6%。消融实验表明,报告的融合轨迹在组件数、局部秩和模态纯度阈值间稳定。这些结果支持局部几何邻域作为分析所评估VLMs跨模态表示的有用可解释和因果单元。
英文摘要
Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modalities interact remains poorly understood. Most interpretability methods identify global linear directions, which may miss representations that are globally high-dimensional but locally low-dimensional. We introduce LENS (Local Explanation of Neighborhood Subspaces), a method that decomposes VLM activations into local low-rank Gaussian neighborhoods using a Mixture of Factor Analyzers. Applied to LLaVA-1.5-7B and Qwen3-VL-8B, LENS reveals distinct depth-dependent fusion trajectories consistent with each model's fusion mechanism: LLaVA progressively mixes modalities at later layers, whereas Qwen3-VL mixes them early, partially re-segregates them, and recombines them near the output. An automated multimodal labeling pipeline assigns concise semantic descriptions to these neighborhoods. Interpolating activations toward neighborhood centroids causally redirects generation within and across modalities and outperforms difference-in-means and VL-SAE in most evaluated conditions; in one LLaVA vision-to-vision setting, MFA achieves 5.7 times the VL-SAE score. Human evaluation finds MFA steering competitive with prompting and substantially stronger than the other intervention baselines. Finally, the MFA coefficient space improves Qwen3-VL image-to-rendered-text retrieval at the deepest evaluated layer from 14.9% to 48.6% R@1. Ablations show that the reported fusion trajectories are stable across component counts, local ranks, and modality-purity thresholds. These results support local geometric neighborhoods as useful interpretable and causal units for analyzing cross-modal representations in the evaluated VLMs.