发表机构
Michigan State University; University of North Carolina at Chapel Hill(密歇根州立大学; 北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对细粒度视觉识别提出SubViT方法,通过多子词表示有区别补丁并结合全局上下文,采用两阶段训练策略。在多个数据集上评估,提升了新类别准确率,增加少量延迟和FLOP,减少了相对于Retina Patch的延迟,证明了更广泛适用性。
AI 中文摘要
我们提出了子词视觉Transformer(SubViT),一种用于细粒度视觉识别的选择性图像令牌化方法。标准视觉Transformer将每个固定大小的补丁压缩为单个令牌,而细粒度差异往往仅取决于少数补丁内的局部变化。SubViT通过用多个子词表示有区别的补丁,同时保留用于全局上下文的原始令牌序列来解决这种不匹配,从而在最需要的地方分配额外的能力。由于注意力头编码互补语义,并且在推理时提取注意力图需要额外的主干前向传播,我们采用两阶段训练策略。阶段1使用从随机注意力头采样的细分区域对ViT进行微调,使模型接触不同的细分模式。阶段2通过特征退化距离识别信息丰富的注意力图,并将其提炼为轻量级单图路由器,该路由器直接预测确定性令牌重要性分数,无需单独的注意力前向传播。我们在广义类别发现(GCD)上评估SubViT,这是一项具有挑战性的任务,需要细粒度辨别和对未标记新类别进行泛化。在CUB、FGVC-Aircraft和Stanford Cars数据集上,SubViT将DINOv2的平均新类别准确率从81.3%提高到84.7%,仅增加0.50毫秒的延迟和3.4%的FLOP,同时相对于Retina Patch减少73.8%的延迟。在CIFAR-10和ImageNet-100上的结果证明了其更广泛的适用性。
英文摘要
We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transformers compress each fixed-size patch into a single token, although fine-grained distinctions often depend on localized variations within only a few patches. SubViT addresses this mismatch by representing discriminative patches with multiple subtokens while retaining the original token sequence for global context, thereby allocating additional capacity where it is most needed. Since attention heads encode complementary semantics and extracting attention maps at inference requires an extra backbone forward, we adopt a two-stage training strategy. Stage 1 fine-tunes the ViT using subdivision regions sampled from random attention heads, exposing the model to diverse subdivision patterns. Stage 2 identifies informative attention maps through feature-degradation distances and distills them into a lightweight single-map router, which directly predicts deterministic token-importance scores without a separate attention forward. We evaluate SubViT on Generalized Category Discovery (GCD), a challenging task requiring both fine-grained discrimination and generalization to unlabeled novel categories. Across CUB, FGVC-Aircraft, and Stanford-Cars, SubViT improves the average novel-category accuracy of DINOv2 from $81.3\%$ to $84.7\%$, with only $0.50$ ms additional latency and $3.4\%$ more FLOPs, while reducing latency by $73.8\%$ relative to Retina Patch. Code: \href{https://github.com/jiezhu23/SubViT_ACCV26}{SubViT}.