HiPerViT:用于多尺度纹理识别的层次化感知器-视觉Transformer架构
HiPerViT: A Hierarchical Perceiver-Vision Transformer Architecture for Multi-Scale Texture Recognition
- Institute of Mathematics and Computer Sciences (ICMC), University of São Paulo (USP)(圣保罗大学数学与计算机科学研究所)
- São Carlos Institute of Physics (IFSC), University of São Paulo (USP)(圣保罗大学圣卡洛斯物理研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
HiPerViT通过将二阶统计先验编码为统计令牌并集成到Transformer流程中,在六个纹理基准上显著提升识别性能,验证了显式统计令牌化的有效性。
AI中文摘要:
纹理识别对现代视觉模型而言仍具挑战性,因为判别性证据通常由高阶空间统计特征而非仅由物体形状承载。尽管视觉Transformer提供了强大的长距离建模能力,但其标准的以对象为中心的表示并未显式暴露此类统计结构,这限制了细粒度识别场景中的纹理敏感性。我们提出HiPerViT,一种紧凑的纯视觉架构,将显式的二阶统计先验注入基于Transformer的识别流程中。该方法结合全局与局部图像视图,将紧凑双线性描述符编码为统计令牌,并通过感知器式潜在蒸馏将该令牌与一阶空间表示集成。该设计使空间令牌与二阶特征共现统计之间能够直接交互,为模型提供对纹理相关信息的显式访问,而无需多模态预训练或集成构建。在六个纹理识别基准上,HiPerViT在报告的评估协议下相较于强纯视觉基线取得了一致改进,包括在DTD上提升+3.05个百分点,在GTOS-Mobile上提升+10.48,在1200Tex上提升+10.10。除基准性能外,我们的分析表明这些提升在很大程度上不受用于提取二阶统计的骨干网络深度以及交互与蒸馏阶段顺序的影响。这一模式表明改进的主要来源并非特定的融合拓扑,而是二阶统计信息作为一等表示信号的显式可用性。这些结果支持将显式统计令牌化作为面向纹理的视觉识别的有效且稳健的设计原则。
英文摘要:
Texture recognition remains challenging for modern vision models because discriminative evidence is often carried by higher-order spatial statistics rather than by object shape alone. While Vision Transformers provide strong long-range modeling capacity, their standard object-centric representations do not explicitly expose such statistical structure, which limits texture sensitivity in fine-grained recognition settings. We present HiPerViT, a compact vision-only architecture that injects an explicit second-order statistical prior into a transformer-based recognition pipeline. The method combines global and local image views with a compact bilinear descriptor encoded as a statistical token, and integrates this token with first-order spatial representations through Perceiver-style latent distillation. This design enables direct interaction between spatial tokens and second-order feature co-occurrence statistics, providing the model with explicit access to texture-relevant information without requiring multimodal pretraining or ensemble construction. Across six texture recognition benchmarks, HiPerViT achieves consistent improvements over strong vision-only baselines under the reported evaluation protocols, including gains of +3.05 percentage points on DTD, +10.48 on GTOS-Mobile, and +10.10 on 1200Tex. Beyond benchmark performance, our analyses show that these gains are largely invariant to the backbone depth used to extract second-order statistics and to the ordering of interaction and distillation stages. This pattern suggests that the primary source of improvement is not a specific fusion topology, but the explicit availability of second-order statistical information as a first-class representational signal. These results support explicit statistical tokenization as an effective and robust design principle for texture-centric visual recognition.