arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16263cs.CV

先见后答:面向视觉语言模型的无训练视觉层分析

Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models

Ruchen Liu, Yi Yang, Yiming Xu, Michael Ying Yang, Monika Sester, Bodo Rosenhahn

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出无需训练的视觉数据集熵(VDE)方法,可预测视觉语言模型的最优视觉层,缩小搜索范围,相比GW距离更具实用性。

中文摘要 AI 辅助

LLaVA风格的视觉语言模型(VLM)将视觉骨干网络的固定后期层(通常是倒数第二层)的视觉令牌传递给语言模型。我们首先证明这一隐藏惯例是脆弱的:在2个VLM和7个图像及视频基准测试中,14个模型-任务对里有13个的默认层并非最优,且最优层会随任务和视觉骨干网络变化。通过逐层穷举推理来寻找该层的成本过高,且不存在更好的固定默认值。因此我们提出,能否从表示几何结构预测层的有用性。我们研究了用于单模态层分析的基于矩阵的熵,将其在样本级视觉嵌入上计算为视觉数据集熵(VDE);还研究了用于编码器级VLM模型选择的格罗莫夫-瓦瑟斯坦(GW)距离,我们将其重新用作逐层视觉-语言对齐信号。将这些方法迁移到基于LLaVA的模型并非先验显而易见:视觉塔是冻结的,而多模态投影器是训练过的,因此我们对投影器两侧都进行分析。我们发现VDE可迁移,而GW不可。从100个无标签任务样本计算,无需下游推理,投影器前的VDE可追踪逐层准确率,其排名最高的层在基于SigLIP的LLaVA-Video的每个任务上都覆盖了最优层,同时为基于CLIP的Video-LLaVA提供了区域级指导。投影器后的分析显示,投影器重塑了视觉几何结构,但未消除与性能相关的趋势,使得VDEpre成为更强的信号。而GW在投影后被扁平化,最好被解读为对齐诊断而非选择器。因此VDE提供了一种可解释的、无训练的策略,可将视觉层搜索范围缩小到少数候选,以进行有限的下游验证。

英文摘要

LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model. We first show that this hidden convention is fragile: across 2 VLMs and 7 image and video benchmarks, the default layer is sub-optimal in 13 of 14 model-task pairs, and the best layer shifts with both task and visual backbone. Finding that layer by exhaustive layer-wise inference is prohibitively expensive, and no better fixed default exists. We therefore ask whether layer usefulness can instead be predicted from representation geometry. We study matrix-based entropy, introduced for unimodal layer analysis, which we compute over sample-level visual embeddings as Visual Dataset Entropy (VDE); and Gromov-Wasserstein (GW) distance, introduced for encoder-level VLM model selection, which we repurpose as a layer-wise visual--language alignment signal. Transferring these to LLaVA-based models is not obvious a priori: the vision tower is frozen while the multimodal projector is trained, so we profile both sides of the projector. We find that VDE transfers, and GW does not. Computed from 100 unlabeled task samples without downstream inference, pre-projector VDE tracks layer-wise accuracy and its top-ranked layers cover the oracle best layer on every task for the SigLIP-based LLaVA-Video, while giving region-level guidance for the CLIP-based Video-LLaVA. Post-projector profiles show that the projector reshapes visual geometry but does not erase the performance-relevant trend, leaving $\mathrm{VDE}_{\mathrm{pre}}$ the stronger signal. GW instead flattens after projection and is best read as an alignment diagnostic rather than a selector. VDE thus offers an interpretable, training-free policy that narrows the visual-layer search to a handful of candidates for limited downstream verification.

发表机构

  • Leibniz Universität Hannover(汉诺威莱布尼茨大学)
  • University of Bath(巴斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑