Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
VLMs 是在看还是只是在说?揭示视觉重新检查的幻觉
机构 * University of Southern California(南加州大学) ; University of California San Diego(加州大学圣地亚哥分校) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);grounding(abstract);分类 cs.CV
AI总结 通过图像交换探测框架 VisualSwap 和 800 对图像基准 VS-Bench,发现视觉语言模型在推理时声称的“重新检查图像”多为文本模式,而非真正的视觉重新检查,且思考模型更易受影响,用户指令可恢复视觉基础但自我反思无效。
Comments ICML 2026 Oral