arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26590cs.CVcs.LG

GTR:用于高效密集预测的门控令牌循环

GTR: Gated Token Recurrence for Efficient Dense Prediction

Zhe Feng, Longfei Liu, Wei Liu, Kai Chen, Jiangang Kong, Wei Zhou, Yifeng Qian, Dexiong Chen, Xuanlong Yu, Xi Shen

首次发表
浏览论文内容

中文总结 AI 辅助

提出GTR,一种无softmax的循环视觉骨干,结合门控线性注意力与空间增强,从DINOv3蒸馏,在COCO上达到58.9 AP,且延迟低,为高分辨率密集预测提供高效替代方案。

中文摘要 AI 辅助

基于自注意力的视觉骨干网络在密集预测任务上表现良好,但全局softmax注意力的二次方计算成本限制了其在图像分辨率增加时的效率。我们引入了门控令牌循环(GTR),一种无softmax的循环视觉骨干网络,它结合了门控线性注意力、交替空间扫描方向和空间增强的SwiGLU块。GTR从检测专用的DINOv3教师模型中蒸馏而来,仅使用最终层补丁令牌对齐,通过线性投影和平方ℓ2损失,不使用掩码令牌预测或中间层监督。在Objects365检测器预训练下,GTR-L在COCO val2017上实现了58.9的框AP,在RTX 4090上编译的FP16执行下,中位批量一延迟为1.908毫秒。同一骨干网络还迁移到实例分割、姿态估计、定向检测、语义分割和单目深度估计。在隔离的内核基准测试中,我们的专用分块CUDA算子比FLA v0.5.0在RTX 4090上处理1.6K令牌时快4.0倍。在DRIVE AGX Thor上的TensorRT部署,在评估的模型中实现了2.282–8.769毫秒的中位批量一延迟。这些结果表明,循环令牌混合可以为高分辨率密集预测提供全局softmax注意力的高效替代方案,并边缘化此http URL页面:此https URL。

英文摘要

Self-attention-based vision backbones perform well on dense prediction, but the quadratic computational cost of global softmax attention limits their efficiency as image resolution increases. We introduce Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone that combines gated linear attention, alternating spatial scan directions, and spatially enhanced SwiGLU blocks. GTR is distilled from a detection-specialized DINOv3 teacher using only final-layer patch-token alignment through a linear projection and squared $\ell_2$ loss, without masked-token prediction or intermediate-layer supervision. With Objects365 detector pre-training, GTR-L achieves 58.9 box AP on COCO \texttt{val2017} with 1.908\,ms median batch-one latency under compiled FP16 execution on an RTX~4090. The same backbone also transfers to instance segmentation, pose estimation, oriented detection, semantic segmentation, and monocular depth estimation. In an isolated kernel benchmark, our specialized chunkwise CUDA operator is $4.0\times$ faster than FLA v0.5.0 at 1.6K tokens on RTX~4090. TensorRT deployment on DRIVE AGX Thor achieves 2.282--8.769\,ms median batch-one latency across the evaluated models. These results show that recurrent token mixing can provide an efficient alternative to global softmax attention for high-resolution dense prediction and edge deployment. Project page: https://intellindust-ai-lab.github.io/projects/GTR/

发表机构

  • Didi International Business Group(滴滴国际业务集团)
  • Intellindust AI Lab(Intellindust AI实验室)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • Didi Research(滴滴研究院)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑