IronViT:迈向高效通用视觉表示学习
IronViT: Toward Efficient Generalist Visual Representation Learning
AI总结:
IronViT通过先蒸馏专家能力至softmax桥再迁移至混合线性注意力编码器,实现了高效且通用的视觉表示学习,在多项任务中媲美专家模型并降低高分辨率计算成本。
AI中文摘要:
通用视觉编码器必须在统一表示中捕获语义、空间、语言对齐和与动作相关的线索,然而当今最强大的视觉骨干网络所依赖的softmax注意力在高分辨率下变得极其昂贵。解决这两个挑战的一种自然尝试是将多个专家教师直接蒸馏到一个高效架构中。我们发现,直接耦合这些目标会降低表示质量,因为学生网络必须同时协调异构能力并将它们适应到不同的令牌混合架构中。我们提出了IronViT,它基于一个简单原则:在约束计算之前整合能力。IronViT首先将互补的专家蒸馏到一个softmax注意力能力桥中,然后逐步将整合后的表示转移到混合softmax-线性注意力编码器。一个专门构建的数据管道进一步策划蒸馏语料库,以提高信息密度和更广泛的领域覆盖。在识别、检索、密集预测、多模态理解和机器人学习方面,IronViT与领先的专家和通用视觉编码器相比具有竞争力。softmax桥在评估的骨干网络中在多模态理解和机器人学习方面实现了最强的综合性能,而混合编码器保留了广泛的迁移性能,其效率优势随输入分辨率增长而增加。这些结果共同表明,在架构转换之前整合能力可以产生一个通用视觉编码器,而无需继承传统softmax注意力的高昂高分辨率成本。
英文摘要:
A generalist vision encoder must capture semantic, spatial, language-aligned, and action-relevant cues within a unified representation, yet softmax attention underlying today's most capable visual backbones becomes prohibitively expensive at high resolution. A natural attempt to address both challenges is to distill multiple specialist teachers directly into an efficient architecture. We find that directly coupling these objectives degrades representation quality, as the student must simultaneously reconcile heterogeneous capabilities and adapt them to a different token-mixing architecture. We introduce IronViT, built on a simple principle: consolidate capabilities before constraining computation. IronViT first distills complementary specialists into a softmax attention capability bridge, then progressively transfers the consolidated representation to a hybrid softmax-linear attention encoder. A purpose-built data pipeline further curates the distillation corpus for higher information density and broader domain coverage. Across recognition, retrieval, dense prediction, multimodal understanding, and robotic learning, IronViT is competitive with leading specialist and generalist vision encoders. The softmax bridge achieves the strongest aggregate performance in multimodal understanding and robotic learning among the evaluated backbones, while the hybrid encoder retains broad transfer performance with an efficiency advantage that grows with input resolution. Together, these results show that consolidating capabilities before architectural conversion can yield a generalist visual encoder without inheriting the prohibitive high-resolution cost of conventional softmax attention.