arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.25178cs.CVcs.AIcs.LG

GHOST:诱导幻觉的多模态大语言模型图像生成

GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs

  • The University of Melbourne(墨尔本大学)
  • University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

Aryan Yazdan Parast, Parsa Hosseini, Hesam Asadollahzadeh, Arshia Soltani Moakhar, Basim Azam, Soheil Feizi, Naveed Akhtar

更新

AI总结:

GHOST通过优化隐蔽令牌生成诱导幻觉的图像,评估多模态大语言模型的可靠性,并发现高幻觉成功率及可转移漏洞。

AI中文摘要:

多模态大语言模型(MLLMs)中的对象幻觉是一种持续存在的故障模式,导致模型感知到图像中不存在的对象。目前,该弱点是通过静态基准和固定视觉场景进行研究的,这预设了发现模型特定或未预料到的幻觉漏洞的可能性。我们引入了GHOST(通过优化隐蔽令牌生成幻觉),一种旨在通过主动生成诱导幻觉的图像来压力测试MLLMs的方法。GHOST完全自动,不需要人工监督或先前知识。它通过在图像嵌入空间中优化以误导模型,同时保持目标对象不存在,然后引导扩散模型根据嵌入生成自然的图像。生成的图像在视觉上自然且接近原始输入,但引入了微妙的误导线索,导致模型产生幻觉。我们评估了我们的方法在各种模型上的表现,包括推理模型如GLM-4.1V-Thinking,并实现了超过28%的幻觉成功率,相比之下,先前数据驱动的方法约为1%。我们通过定量指标和人工评估确认生成的图像既高质量又无对象。此外,GHOST揭示了可转移的漏洞:针对Qwen2.5-VL优化的图像在GPT-4o上导致66.5%的幻觉率。最后,我们展示了在我们的图像上微调可以缓解幻觉,使GHOST成为构建更可靠多模态系统诊断和纠正工具。

英文摘要:

Object hallucination in Multimodal Large Language Models (MLLMs) is a persistent failure mode that causes the model to perceive objects absent in the image. This weakness of MLLMs is currently studied using static benchmarks with fixed visual scenarios, which preempts the possibility of uncovering model-specific or unanticipated hallucination vulnerabilities. We introduce GHOST (Generating Hallucinations via Optimizing Stealth Tokens), a method designed to stress-test MLLMs by actively generating images that induce hallucination. GHOST is fully automatic and requires no human supervision or prior knowledge. It operates by optimizing in the image embedding space to mislead the model while keeping the target object absent, and then guiding a diffusion model conditioned on the embedding to generate natural-looking images. The resulting images remain visually natural and close to the original input, yet introduce subtle misleading cues that cause the model to hallucinate. We evaluate our method across a range of models, including reasoning models like GLM-4.1V-Thinking, and achieve a hallucination success rate exceeding 28%, compared to around 1% in prior data-driven discovery methods. We confirm that the generated images are both high-quality and object-free through quantitative metrics and human evaluation. Also, GHOST uncovers transferable vulnerabilities: images optimized for Qwen2.5-VL induce hallucinations in GPT-4o at a 66.5% rate. Finally, we show that fine-tuning on our images mitigates hallucination, positioning GHOST as both a diagnostic and corrective tool for building more reliable multimodal systems.

↑