发表机构
MBZUAI; Imperial College London(穆罕默德·本·扎耶德人工智能大学; 伦敦帝国学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型计数失败问题,通过对VLM激活进行探测训练及分析,发现其虽常编码正确计数但输出错误,提出探测器引导的自校正方法,在不更新参数时提高计数准确率,揭示内部知识与模型输出差距并提供改进工具。
AI 中文摘要
尽管视觉语言模型(VLM)在许多多模态任务中表现出色,但在基本物体计数方面仍存在困难。研究其是缺少内部知识还是内部表征与语言输出之间存在差距。在五个计数数据集上对四个VLM的激活进行简单探测训练,发现非线性探测可检测计数错误,VLM常编码正确计数却输出错误答案。SVCCA分析表明相关探测占据部分共享激活子空间但读出方向未对齐。通过因果转向干预验证结果,提出探测器引导的自校正方法,在不更新参数情况下提高计数准确率达15.6个绝对百分点,为改进VLM计数提供实用工具并揭示内部知识与模型输出差距。
英文摘要
Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting. We investigate whether this reflects missing internal knowledge or a gap between internal representations and verbalized outputs. Training simple probes on activations from four VLMs across five counting datasets reveals that nonlinear probes can reliably detect counting errors, suggesting that VLMs often encode the correct count even when they output the wrong answer. SVCCA analysis shows that probes trained on ground-truth counts and probes trained on model outputs occupy a partially shared activation subspace but read out along misaligned directions. We further validate our findings using a causal steering intervention, proving that strengthening the direction of count-identified probes does improve model counting performance. Motivated by this result, we propose a detector-guided self-correction method that selectively re-prompts the model only when an internal error detector predicts failure. This simple inference-time intervention improves counting accuracy by up to 15.6 absolute percentage points, without any parameter updates. Our results establish activation-based error probing as both a practical tool for improving VLM counting and a mechanistic lens on the gap between internal knowledge and model outputs.