发表机构
King Abdullah University of Science and Technology (KAUST); Harbin Institute of Technology, Shenzhen(阿卜杜拉国王科技大学; 哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在CLIP的最终视觉自注意力层用α-entmax变换替代逐行softmax,以解决其在密集开放词汇预测时注意力分散产生噪声的问题,在开放词汇任务评估中,注意力稀疏化增益与基线注意力偏离目标类程度成正比。
AI 中文摘要
对比语言-图像预训练(CLIP)依赖基于softmax的自注意力,这是一种严格的正分布,会将概率质量分配给每对token,包括语义无关的。虽然这些密集的softmax权重在预训练期间收集广泛上下文有效,但会在许多低显著性token上分散注意力,产生噪声。我们研究在最终视觉自注意力层中用α-entmax变换替代逐行softmax,应用于标准查询-键注意力和自相关变体。因为entmax应用数据依赖阈值将低分数精确映射到零,它起到隐式去噪器作用,将上下文无关依赖归零,同时将质量重新分配到最相关token上。我们在开放词汇任务上评估,发现注意力稀疏化的增益与基线注意力偏离目标类的程度成正比。
英文摘要
Contrastive Language-Image Pre-training (CLIP) relies on softmax-based self-attention, a strictly positive distribution that assigns probability mass to every pair of tokens-even semantically irrelevant ones. While these dense softmax weights are effective for gathering broad context during pre-training, they spread attention across many low-salience tokens, producing noise that obscures the fine-grained, spatially localized cues required for dense, open-vocabulary prediction. We study an inference-time substitution of the row-wise softmax in the final visual self-attention layers with the $α$-entmax transform, applied across both the standard query-key attention and self-correlation variants. Because entmax applies a data-dependent threshold that maps low scores exactly to zero, it acts as an implicit denoiser, zeroing contextually irrelevant dependencies while redistributing mass onto the most relevant tokens. We evaluate on open-vocabulary tasks-dense semantic segmentation (Pascal VOC, Pascal Context, ADE20K) and fine-grained retrieval (FG-OVD)-and find the gain from attention sparsification is proportional to how much the baseline attention spreads off the target class.