arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当视觉信号产生误导:视觉语言模型中属性幻觉的机制研究

When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

Yufei Zhang, Chenlu Zhan, Hongwei Wang

arXiv 2608.11024首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉语言模型的属性幻觉问题,提出VISOR框架,通过VSNR诊断区分失败模式并路由适配操作,在三类模型上减少属性假阳性且不依赖先验主导假设。

AI 中文摘要

属性幻觉指视觉语言模型(VLMs)能正确识别物体却错误描述其属性,该现象普遍存在但机制研究不足。主流解释为语言先验主导,这推动了先验抑制方法,但该解释未在属性层面直接验证。本文提出VISOR(Visual-Operational Remediation),即统一框架,结合空图像诊断与路由修复。其VSNR诊断将每个预测分解为视觉logit信号与语言先验信号。在来自三类VLMs家族、三类属性类型的10791个负真值样本中,视觉信号可强预测假阳性,而语言先验信号接近随机水平。VISOR利用该诊断区分两种失败模式:颜色/状态属性中低边际但方向正确的视觉信号,以及材质属性中低SNR或未对齐的视觉信号。同一诊断将每个查询路由至合适操作:针对阈值设置错误的校准、针对无训练低SNR处理的弃权(不执行),或针对先验抑制无法纠正的材质故障的定向视觉适配。在Qwen、InternVL和LLaVA上,VISOR减少了属性假阳性,且不依赖先验主导假设。

英文摘要

Attribute hallucination---where vision-language models (VLMs) correctly identify an object but mischaracterize its properties---is prevalent yet mechanistically poorly understood. The dominant explanation, language-prior dominance, has motivated prior-suppression methods, but this explanation has not been directly tested at the attribute level. We present VISOR (Visual-Operational Remediation), a unified framework that couples null-image-based diagnosis with routed remediation. Its VSNR diagnostic decomposes each prediction into a visual logit signal and a language-prior signal. Across 10,791 negative-ground-truth samples from three VLM families and three attribute types, the visual signal strongly predicts false positives, whereas the language-prior signal is near chance. VISOR uses this diagnosis to separate two failure modes: low-margin but directionally correct visual signals in color/state attributes, and low-SNR or misaligned visual signals in material attributes. The same diagnosis routes each query to the appropriate operator: calibration for threshold-placement errors, abstention for training-free low-SNR handling, or targeted visual adaptation for material failures that prior suppression cannot correct. Across Qwen, InternVL, and LLaVA, VISOR reduces attribute false positives without relying on the prior-dominance assumption.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑