发表机构
University of California, Los Angeles; Peking University(加州大学洛杉矶分校; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出利用MoE视觉语言模型内部路由概率检测目标缺失并引导选择性视觉重定位,无需修改权重,显著提升端到端准确率。
AI 中文摘要
视觉语言模型(VLMs)可能接受错误的视觉前提,即使目标对象不存在,也会回答关于其颜色、数量、位置或状态的问题。我们将这种可靠性关键行为称为目标缺失接地失败。现有的视觉接地检测器主要依赖于生成的响应、隐藏状态或不确定性度量。我们提出了第一个利用混合专家(MoE)视觉语言模型中的内部路由决策来在生成前检测目标缺失并指导选择性纠正的框架。我们从Qwen3-VL-30B-A3B-Instruct和Gemma-4-26B-A4B-it中提取目标令牌路由概率,为每个模型训练一个独立的L2正则化线性检测器,并利用其预测选择性地调用目标感知审查提示。仅使用路由,Qwen和Gemma检测器在GQA-Inpaint上分别达到0.9988和0.9956的ROC-AUC,在外部OBER数据集上分别保持0.8095和0.7781。由此产生的路由门控策略在GQA-Inpaint和OBER上分别将端到端准确率提高了Qwen的+22.25%和+12.17%,以及Gemma的+13.42%和+1.39%,且无需修改模型权重。进一步分析表明,该信号定位于目标对象令牌,在MoE早期层中出现,并分布在部分可替代的专家中。尽管跨数据集阈值偏移需要重新校准,但误报审查总体造成的损害有限,这表明可以通过联合选择检测器阈值和审查提示来控制干预风险。总体而言,我们表明仅路由概率就保留了关于视觉感知的可操作信息,使得MoE VLM已产生的计算能够支持低成本检测和选择性视觉重定位。
英文摘要
Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, hidden states, or uncertainty measures. We present the first framework to leverage internal routing decisions in Mixture-of-Experts (MoE) VLMs to detect target absence before generation and guide selective correction. We extract target-token routing probabilities from Qwen3-VL-30B-A3B-Instruct and Gemma-4-26B-A4B-it, train a separate L2-regularized linear detector for each model, and use its predictions to selectively invoke a target-aware review prompt. Using routing alone, the Qwen and Gemma detectors achieve ROC-AUCs of 0.9988 and 0.9956 on GQA-Inpaint and retain 0.8095 and 0.7781 on the external OBER dataset, respectively. The resulting routing-gated policy improves end-to-end accuracy on GQA-Inpaint and OBER by +22.25% and +12.17% for Qwen, and by +13.42% and +1.39% for Gemma, without modifying model weights. Further analysis shows that the signal is localized to the target-object token, emerges in early MoE layers, and is distributed across partially substitutable experts. Although cross-dataset threshold shifts require recalibration, false-positive review causes limited harm overall, suggesting that intervention risk can be controlled through joint selection of the detector threshold and review prompt. Overall, we show that routing probabilities alone preserve actionable information about visual perception, allowing computation already produced by an MoE VLM to support low-cost detection and selective visual regrounding.