arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGE:面向视觉-语言解码器中视觉定位的感知汇聚引导增强

SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders

Jeonghyo Song, YoungJoon Yoo

arXiv 2610.11469首次发表:更新:

发表机构

Chung-Ang University(中央大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SAGE是一种轻量级干预方法,通过引导视觉-语言解码器注意力远离提示不变汇聚点,提升视觉定位性能、减少幻觉,在多类VLM基准中取得一致增益。

AI 中文摘要

近期的大型视觉-语言模型(VLMs)将视觉编码器与大型语言模型(LLM)相结合,在各类图像-文本任务中表现出色,但其可靠性常受解码器注意力异常的限制,这些异常会抑制视觉证据并加剧幻觉。本文重新审视了视觉注意力汇聚点,发现了一种结构化的、依赖层的行为:在不同提示下,解码器的早期和晚期层会出现对相同少数图像区域的提示不变注意力崩溃,我们将其称为PIS(提示不变汇聚点);而中间层则会根据提示调整,驱动视觉-语言对齐。这种划分表明,将汇聚点视为单一效应的处理方式并不完整。基于这一见解,我们提出了SAGE(Sink-Aware Guided Emphasis,感知汇聚引导增强),这是一种轻量级干预方法,利用来自CLIP、ViT、DINOv3等标准视觉骨干的与令牌对齐的感兴趣区域(ROI)掩码,引导解码器注意力远离PIS,转向查询依赖的感兴趣区域。在各类视觉编码器+仅解码器LLM的VLM家族上进行评估,当使用骨干派生的ROI掩码实例化时,SAGE可提升视觉定位性能、减少幻觉,并在公共下游视觉-语言基准(包括需要局部证据的细粒度视觉判别场景)中取得一致增益。

英文摘要

Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations. In this paper, we revisit visual attention sinks and uncover a structured, layer-dependent behavior: across prompts, early and late decoder layers exhibit prompt-invariant attention collapse onto the same few image regions, which we term PIS (Prompt-Invariant Sinks), whereas mid layers become prompt-conditioned and drive vision-language alignment. This split suggests that treating sinks as a uniform effect is incomplete. Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3. Evaluated on diverse vision-encoder + decoder-only LLM VLM families, SAGE improves visual grounding, reduces hallucinations, and yields consistent gains across public downstream vision-language benchmarks, including fine-grained visual discrimination settings where localized evidence is crucial, when instantiated with backbone-derived ROI masks.

CommentsAccepted to EMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑