GTR:用于高效密集预测的门控令牌循环
GTR: Gated Token Recurrence for Efficient Dense Prediction
浏览论文内容
中文总结 AI 辅助
提出GTR,一种无softmax的循环视觉骨干,结合门控线性注意力与空间增强,从DINOv3蒸馏,在COCO上达到58.9 AP,且延迟低,为高分辨率密集预测提供高效替代方案。
中文摘要 AI 辅助
基于自注意力的视觉骨干网络在密集预测任务上表现良好,但全局softmax注意力的二次方计算成本限制了其在图像分辨率增加时的效率。我们引入了门控令牌循环(GTR),一种无softmax的循环视觉骨干网络,它结合了门控线性注意力、交替空间扫描方向和空间增强的SwiGLU块。GTR从检测专用的DINOv3教师模型中蒸馏而来,仅使用最终层补丁令牌对齐,通过线性投影和平方ℓ2损失,不使用掩码令牌预测或中间层监督。在Objects365检测器预训练下,GTR-L在COCO val2017上实现了58.9的框AP,在RTX 4090上编译的FP16执行下,中位批量一延迟为1.908毫秒。同一骨干网络还迁移到实例分割、姿态估计、定向检测、语义分割和单目深度估计。在隔离的内核基准测试中,我们的专用分块CUDA算子比FLA v0.5.0在RTX 4090上处理1.6K令牌时快4.0倍。在DRIVE AGX Thor上的TensorRT部署,在评估的模型中实现了2.282–8.769毫秒的中位批量一延迟。这些结果表明,循环令牌混合可以为高分辨率密集预测提供全局softmax注意力的高效替代方案,并边缘化此http URL页面:此https URL。
英文摘要
Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated linear attention, alternating spatial scan directions, and spatially enhanced SwiGLU blocks. GTR is distilled from a detection-specialized DINOv3 teacher using only final-layer patch-token alignment through a linear projection and squared $\ell_2$ loss, without masked-token prediction or intermediate-layer supervision. With Objects365 detector pre-training, GTR-L achieves 58.9 box AP on COCO \texttt{val2017} with 1.908\,ms median batch-one latency under compiled FP16 execution on an RTX~4090. The same backbone also transfers to instance segmentation, pose estimation, oriented detection, semantic segmentation, and monocular depth estimation. In an isolated kernel benchmark, our specialized chunkwise CUDA operator is $4.0\times$ faster than FLA v0.5.0 at 1.6K tokens on RTX~4090. TensorRT deployment on DRIVE AGX Thor achieves 2.282--8.769\,ms median batch-one latency across the evaluated models. These results show that recurrent token mixing can provide an efficient alternative to global softmax attention for high-resolution dense prediction and edge deployment. Project page: https://intellindust-ai-lab.github.io/projects/GTR/
发表机构
- Didi International Business Group(滴滴国际业务集团)
- Intellindust AI Lab(Intellindust AI实验室)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- Didi Research(滴滴研究院)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。