arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在注意力流形中分离语义注意力与结构偏差

Disentangling Semantic Attention from Structural Bias in the Attention Manifold

Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-gang Jiang

arXiv 2607.24017首次发表:更新:

发表机构

Fudan University; Singapore Management University(复旦大学; 新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态大语言模型中注意力机制对无信息视觉 tokens 关注过多的问题,提出无需训练的 SPAR 方法,通过净化噪声和重分配注意力,有效恢复真实视觉基础且计算开销小。

AI 中文摘要

多模态大语言模型(MLLMs)中注意力机制的经验成功往往掩盖了其内在的细微缺陷。MLLMs 始终对某些语义上无信息的视觉 tokens 表现出不成比例的关注,即 “注册” 或 “视觉注意力汇” 现象。现有推理干预方法孤立处理这些 tokens 且计算效率低。我们将此现象重新构建为对视觉特征施加的广义文本偏差。为此引入 Saliency-guided Purification and Adaptive Redistribution(SPAR),一种无需训练的即插即用干预方法。它通过净化结构噪声减轻偏差,并将回收的注意力预算重新分配到信息最丰富的视觉区域。跨多种幻觉基准的综合评估表明,SPAR 能有效恢复真实视觉基础且计算开销可忽略不计。

英文摘要

The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed "register" or "Visual Attention Sinks." While existing inference intervention methods attempt to identify these sink tokens and redistribute their attention weights, such approaches typically treat these tokens in isolation and suffer from computational inefficiency. Instead, we reframe this phenomenon as a generalized textual bias exerted over visual features that extends beyond isolated sink tokens. From this perspective, a pervasive structural bias leads to the dilution of the semantic visual signal, precipitating multimodal hallucinations as the model prioritizes linguistic priors over valid visual evidence. To address this limitation, we introduce Saliency-guided Purification and Adaptive Redistribution (SPAR), a training-free, plug-and-play intervention. SPAR mitigates this generalized textual bias by purifying structural noise and subsequently redistributing the reclaimed attention budget to the most informative visual regions. Comprehensive evaluations across a diverse spectrum of hallucination benchmarks demonstrate that SPAR effectively restores authentic visual grounding with negligible computational overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑