发表机构
Tel Aviv University(特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种无需训练的事后维纳表示滤波技术,通过离线校准校正视觉语言模型深层前馈输出投影,可在保持运行速度的同时降低其对象幻觉,在多类模型及基准上均验证了有效性与通用性。
AI 中文摘要
视觉语言模型(VLMs)擅长开放式图像描述和视觉问答,但常描述图像中不存在的对象、属性或关系,这一现象被称为对象幻觉。我们提出一种无需训练的事后表示编辑技术,该技术在语言骨干网络的表示空间中运行。该方法在中等规模的配对数据集上执行轻量级、一次性的离线校准,以估计所需的协方差结构,仅使用前向传播和经验二阶统计量,无需梯度更新或微调,之后校正项直接融入模型现有权重。通过将隐藏状态建模为真实成分与幻觉相关成分的叠加,我们推导得到维纳型估计器,其最优增益由配对的真实表示与幻觉表示的协方差以闭式形式给出。特征分解产生满足稳定性准则的逐模态衰减,即滤波器对估计噪声的响应连续。在校正应用于选定深层的前馈输出投影后,推理时模型运行方式不变且速度相同。在LLaVA-1.5、MiniGPT-4、Gemma3和mPLUG-Owl2上的实验表明,该方法在CHAIR、POPE和MME指标上持续降低对象幻觉,同时保持描述流畅性和整体响应质量。我们进一步在TempCompass视频理解基准和用于接地对话的离散扩散语言模型上验证了方法的通用性,结果显示表示滤波即使在时序视频推理和多步骤、全序列去噪场景中也能减少幻觉。
英文摘要
Vision-language models (VLMs) excel at open-ended captioning and visual QA but often describe objects, attributes, or relations absent from the image, a phenomenon known as object hallucination. We propose a {training-free, post-hoc representation editing technique} that operates in the representation space of the language backbone. The method performs a lightweight, one-time offline calibration on a modest paired dataset to estimate the required covariance structures, using only forward passes and empirical second-order statistics with no gradient updates or fine-tuning, after which the correction is absorbed directly into the model's existing weights. By modeling hidden states as a superposition of truthful and hallucination-associated components, we derive a Wiener-type estimator whose optimal gains are given in closed form from the covariances of paired truthful and hallucinated representations. An eigendecomposition yields mode-wise attenuation that respects a stability criterion, i.e., the filter responds continuously to estimation noise. The correction is applied once to the feed-forward output projections of selected deeper layers, at inference time, the model runs unchanged and at the same speed. Experiments on LLaVA-1.5, MiniGPT-4, Gemma3, and mPLUG-Owl2 demonstrate consistent reductions in object hallucination on CHAIR, POPE, and MME while maintaining caption fluency and overall response quality. We further demonstrate the generality of our approach on the TempCompass video understanding benchmark and on discrete diffusion language models for grounded dialogue, showing that representation filtering reduces hallucinations even in temporal video reasoning and multi-step, sequence-wide denoising settings.