发表机构
KC Machine Learning Lab; NFOCZ Inc; Seoul National University(KC机器学习实验室; NFOCZ公司; 首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于窗口群卷积自注意力和分层架构的可扩展旋转反射群等变视觉Transformer,可扩展到百万参数和ImageNet规模数据集,并提供代码和预训练权重。
AI 中文摘要
我们提出了一种可扩展的旋转反射群等变视觉Transformer,其基于窗口群卷积自注意力机制和分层特征架构。我们证明了该方法可扩展到具有数百万参数和实际尺寸图像(即ImageNet)的大型数据集的群等变视觉Transformer(ViT)。所提出的分层窗口旋转反射等变视觉Transformer(REViT-v2)的代码和预训练权重可在该https URL获取。
英文摘要
We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at https://github.com/kc-ml2/revit.
Comments7 pages, Accepted for presentation at NeurIPS NeurREPS workshop 2026