arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小视觉语言模型知道自己何时出错但无法表达:在现实图像退化下对陈述置信度与内部置信度的双模型研究

Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation

M M Asif Ferdous

arXiv 2607.22034首次发表:更新:

AI 中文总结

研究小型视觉语言模型在现实图像退化下陈述置信度与内部置信度的差异,通过评估两个模型在六种退化情况及三种严重程度下的表现,发现内部概率是更好的延迟信号,严重低光时两种信号都不可信。

AI 中文摘要

视觉语言模型(VLM)越来越多地部署在消费硬件上,在这些设备中输入图像会因压缩、相机抖动和光线不足而退化。在这种情况下,可靠的不确定性信号比原始准确率更重要,因为它决定了系统何时应延迟回答而非直接给出答案。我们评估了两个小型开放权重的VLM——Qwen2-VL-2B-Instruct和SmolVLM-Instruct,针对三种严重程度的六种现实照片退化情况,比较了两种置信度信号:模型用自然语言陈述的置信度,以及模型对其生成答案的平均token概率。在3800个预测中,我们发现了一个巨大且一致的差距。Qwen2-VL的语言表达置信度几乎恒定(所有条件下均值为0.87 - 0.90),在随机水平检测自身错误(AUROC为0.39 - 0.75,通常约为0.50),而同一模型的内部token概率以AUROC 0.92 - 0.99区分正确与错误答案。在SmolVLM中,语言表达置信度在很大程度上难以获得:在三个提示模板中,五次试点尝试中只有一次产生了可解析的置信度值,而内部概率再次产生高于随机水平的错误检测(AUROC为0.54 - 0.92)。两个模型在同一情况下失败:在严重曝光不足时,准确率大幅下降(Qwen2-VL从0.99降至0.22,SmolVLM从0.97降至0.42),而两种置信度信号几乎不变,内部错误检测降至随机水平。我们得出结论,小型VLM编码了可用的自我知识,但它们的语言输出并未表达出来,因此在受限部署中内部概率是更好的延迟信号,并且在严重低光条件下两种信号都不应被信任。

英文摘要

Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor lighting. In such settings, a reliable uncertainty signal matters more than raw accuracy, because it determines when a system should defer rather than answer. We evaluate two small open-weight VLMs -- Qwen2-VL-2B-Instruct and SmolVLM-Instruct -- across six realistic photographic degradations at three severity levels, comparing two confidence signals: the confidence the model states in natural language, and the model's own mean token probability over its generated answer. Across 3,800 predictions, we find a large and consistent gap. Verbalized confidence in Qwen2-VL is almost constant (mean 0.87-0.90 across all conditions) and detects its own errors at chance level (AUROC 0.39-0.75, typically ~0.50), while internal token probability from the same model separates correct from incorrect answers with AUROC 0.92-0.99. In SmolVLM, verbalized confidence proved largely unobtainable: across three prompt templates, only one of five pilot attempts produced a parseable confidence value, while internal probability again yielded above-chance error detection (AUROC 0.54-0.92). Both models fail in the same place: under severe underexposure, accuracy collapses (0.99->0.22 for Qwen2-VL, 0.97->0.42 for SmolVLM) while both confidence signals barely move, and internal error-detection falls to chance. We conclude that small VLMs encode usable self-knowledge that their verbalized output does not express, that internal probability is therefore the better deferral signal in constrained deployment, and that neither signal should be trusted under severe low-light conditions.

Comments15 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑