发表机构
Meijo University(名城大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视觉Transformer难以独立评估图像块相关性的问题,提出VisionScreen模型,将筛选机制扩展到视觉识别,通过二维空间域绝对相关性估计,让块选择性聚合相关块,实验证明该方法优于传统ViT,为视觉识别提供新方案。
AI 中文摘要
视觉Transformer(ViT)被广泛用作建模图像块间全局依赖关系的强大框架。但其核心组件自注意力给所有块分配softmax归一化的相对权重,难以独立评估块间相关性。在视觉识别中,图像常含背景或冗余块,自注意力无法明确排除无关块,会引入不必要信息。语言建模领域提出筛选,基于查询-键相似性独立评估每个令牌相关性并通过阈值排除低相关性令牌。本文提出VisionScreen,将筛选机制扩展到视觉识别。它将图像块视为二维网格上的令牌,将基于查询-键相似性的绝对相关性估计扩展到二维空间域。实验表明该方法优于传统ViT,说明筛选对视觉识别有效,为基于softmax注意力的相对特征聚合提供了替代方案。
英文摘要
Vision Transformer (ViT) has been widely used as a powerful framework for modeling global dependencies among image patches. However, its core component, self-attention assigns softmax-normalized relative weights to all patches, making it difficult to evaluate the relevance between patches independently. In visual recognition, images often contain many background or redundant patches, yet self-attention cannot explicitly reject such irrelevant patches, which may introduce unnecessary information into feature aggregation. To address this limitation, Screening has been proposed in the field of language modeling, where the relevance of each token is independently evaluated based on query-key similarity and low-relevance tokens are explicitly excluded through thresholding. In this work, we propose VisionScreen, a new vision model that extends Screening mechanism to visual recognition. VisionScreen treats image patches as tokens arranged on a two-dimensional grid and extends absolute relevance estimation based on query-key similarity to the two-dimensional spatial domain. This allows each patch to selectively aggregate only content-wise and spatially relevant patches without relying on competition among patches. Experiments on image classification benchmarks demonstrate that the proposed method outperforms conventional ViT. These results suggest that Screening can be effective for visual recognition, offering an alternative to relative feature aggregation based on softmax attention.
CommentsExploratory research