发表机构
Zhejiang University; Tencent Hunyuan; Zhejiang Sci-Tech University(浙江大学; 腾讯混元; 浙江理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出基于双曲层次聚类的透明令牌混合器ClusterMixer,构建新骨干HCFormer,其在多视觉任务上性能优于同类方法,推动可解释骨干网络发展。
AI 中文摘要
我们通过重新审视机器学习中最经典的方法之一——聚类,来研究视觉骨干网络中的令牌混合器(token mixer)。有效的令牌混合器是现代视觉骨干网络(如视觉Transformer)的基本组件,可促进图像补丁间的信息交换。主流令牌混合器依赖卷积、注意力、MLP或其混合方法,主要关注准确性与计算成本间的权衡,但存在显著缺陷:其编码过程不透明,缺乏可解释性。与这些不透明设计不同,我们提出ClusterMixer,这是一种基于聚类范式、设计上可解释的透明令牌混合器。ClusterMixer通过层次聚类机制明确构建令牌混合过程;为建模视觉数据固有的类树自然关系,我们在双曲空间中执行聚类,该空间适合以低失真嵌入层次结构。基于此创新,我们提出HCFormer,这是一种将ClusterMixer与一系列精心设计的聚类策略相结合的新骨干架构,以确保在各类任务上的稳健性能。大量实验表明,HCFormer在图像分类、目标检测、实例分割和语义分割等不同任务上始终优于同类方法。鉴于其透明性和有效性,我们希望HCFormer能推动向可解释骨干网络的范式转变。
英文摘要
We investigate the token mixer in vision backbones by revisiting clustering, one of the most classic approaches in machine learning. An effective token mixer is a fundamental component of modern vision backbones like vision Transformers, facilitating information exchange between image patches. Mainstream token mixers, which rely on convolution, attention, MLP, or their hybrids, primarily focus on navigating the trade-off between accuracy and computational cost. However, a significant drawback of these methods is their black-box nature; their encoding process is opaque and lacks interpretability. Diverging from these opaque designs, we introduce ClusterMixer, a transparent token mixer that is grounded in a clustering paradigm and interpretable by design. ClusterMixer explicitly formulates the token mixing process through a hierarchical clustering mechanism. To model the natural, tree-like relationships inherent in visual data, the clustering is performed in hyperbolic space, which is well-suited for embedding hierarchies with low distortion. Building on this innovation, we present HCFormer, a new backbone architecture that integrates ClusterMixer with a series of meticulously designed clustering strategies to ensure robust performance across tasks. Extensive experiments demonstrate that HCFormer consistently outperforms its counterparts across diverse tasks, including image classification, object detection, instance segmentation, and semantic segmentation. Considering its transparency and efficacy, we hope HCFormer can facilitate a paradigm shift toward interpretable backbones.