发表机构
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LVLM物体幻觉,提出免训练的ROT框架,通过检测中间层隐藏状态的语义偏差并施加保范旋转,将其引导回多模态上下文平面,在多个基准上显著降低幻觉。
AI 中文摘要
大型视觉-语言模型(LVLMs)经常遭受物体幻觉问题。现有的免训练干预措施主要操纵注意力权重,这间接影响到达最终预测层的深层语义。在本工作中,我们将焦点转向从自注意力与残差相加之后提取的隐藏状态向量。实证分析揭示,产生幻觉的标记并非简单地过度依赖语言先验;相反,它们表现出异常的上下文偏差,在中间层与文本和视觉上下文的相似性均显著较低。受此启发,我们提出ROT,一个逐层、免训练的框架。ROT在中间层动态检测语义偏差,并应用保范旋转将隐藏状态引导回由上下文张成的局部多模态上下文平面。对于后续层,引入表示平滑机制以稳定校准轨迹。在多个基准上的大量实验表明,ROT在不同模型架构和规模上持续减少幻觉,为接地生成提供了一种高效、几何驱动的解决方案。
英文摘要
Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shift our focus to the hidden state vectors extracted after self-attention and residual addition. Empirical analysis reveals that hallucinated tokens do not simply over-rely on linguistic priors; instead, they exhibit an anomalous contextual deviation, showing significantly lower similarities to both textual and visual contexts in intermediate layers. Motivated by this, we propose ROT, a layer-specific, training-free framework. ROT dynamically detects semantic deviation in the middle layers and applies a norm-preserving rotation to steer the hidden states back toward the local multimodal context plane spanned by the contexts. For subsequent layers, a representational smoothing mechanism is introduced to stabilize the calibrated trajectory. Extensive experiments on multiple benchmarks demonstrate that ROT consistently reduces hallucinations across various model architectures and scales, offering an efficient, geometry-driven solution for grounded generation.
CommentsAccepted in EMNLP 2026 Oral