AI 中文总结
本研究提出reViT,以循环Transformer块结合深度编程专家实现视觉任务,在少参数、低存储下达到高准确率,支持多深度部署与跨任务迁移。
AI 中文摘要
在本研究中,我们表明,以循环方式应用的单个Transformer块,在无需中间特征蒸馏的情况下,可在相当的推理FLOPs下达到全深度视觉编码器的准确率。reViT通过将每个循环深度的FFN表示为小型共享专家库的凸组合,恢复特定深度的变换。连续归一化深度坐标对该混合进行编程,定义了FFN参数空间中可重采样的轨迹。我们在两种场景下评估该设计:监督式ImageNet-1k训练和从DINOv2教师模型进行蒸馏。在两种场景下,受控适配确定,在匹配的单FFN预算下,权重空间合并是测试的最强MoE家族,优于token调度和输出混合替代方案。从头训练的reViT-B/16达到了DeiT III的准确率,存储参数减少约70%。仅使用教师模型的输出特征进行蒸馏的8专家模型,保留了其DINOv2教师模型几乎所有的线性探测准确率,并在分类、分割和深度预测任务中实现迁移。弹性深度训练允许一个检查点(训练好的模型)通过重采样相同的归一化坐标区间,在多个测试深度下运行。对于固定深度部署,循环块可实现为传统的密集图,在不改变每深度单FFN计算的情况下,消除在线路由和合并,同时扩展部署存储。
英文摘要
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70\% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.