arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11826cs.CVcs.AI

从抑制到修复:通过局部分布对齐缓解大型视觉语言模型中的对象幻觉

From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment

  • School of Computer Science, Wuhan University(武汉大学计算机学院)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

Chen Zhao, Xingping Dong, Jiachun Shi, Liang Peng, Chong Wang, Zhen Lei, Ran He, Bo Du

AI总结:

本文提出无需训练的ResOT方法,通过局部分布对齐修复大型视觉语言模型的表示,可减少对象幻觉并提升多模态性能,相关代码将公开。

AI中文摘要:

对象幻觉仍是大型视觉语言模型(LVLMs)生成可靠内容的主要障碍。直观的缓解策略是抑制隐藏表示中与幻觉相关的组件,但这些组件可能包含有用信息,抑制它们会削弱模型的多模态能力。本文提出ResOT,一种无需训练的方法,在推理时通过局部分布对齐修复表示。具体而言,ResOT将主导的幻觉方向从忠实子空间中投影出去,形成用于干预的低维残差子空间;在该子空间内,ResOT使用高斯最优传输(OT)将幻觉分布与忠实分布对齐,得到的映射定义了对原始表示改动最小的修复目标。推理时,ResOT自适应控制每个标记状态向其OT目标移动的距离。在三个代表性LVLMs上的实验表明,ResOT可大幅减少对象幻觉,同时在多个基准测试中提升图像字幕质量和多模态性能,代码将公开。

英文摘要:

Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content. An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations. However, these components may also contain useful information, and suppressing them can weaken the model's multimodal capabilities. In this paper, we propose ResOT, a training-free method that repairs representations at inference time through localized distribution alignment. Specifically, ResOT projects dominant hallucinated directions away from the faithful subspace, forming a low-dimensional residual subspace for intervention. Within this subspace, ResOT uses Gaussian optimal transport (OT) to align the hallucinated distribution with the faithful one. The resulting map defines repair targets with minimal changes to the original representations. At inference, ResOT adaptively controls how far each token state moves toward its OT target. Experiments on three representative LVLMs show that ResOT substantially reduces object hallucination while improving image caption quality and multimodal performance across multiple benchmarks. Code will be released.

↑