arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

硬件感知的校准聚类注意力用于高效视觉几何变换器

Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers

Weitian Wang, Shubham Rai, Cecilia De La Parra, Akash Kumar

arXiv 2610.09274首次发表:更新:

发表机构

Robert Bosch GmbH; Ruhr University Bochum(罗伯特·博世有限公司; 波鸿鲁尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出硬件感知的校准聚类注意力(BC注意力),通过块内聚类和校准方法加速VGGT全局注意力,在GPU上实现2.10-2.63倍加速且性能损失极小。

AI 中文摘要

视觉几何基础变换器(VGGT)在3D场景重建方面标志着重大飞跃,因为它是首个在一次前向传播中直接联合推断所有关键3D属性(相机位姿、深度和稠密几何)的模型。然而,这种联合推断机制需要具有极长序列的全局注意力层,这导致了显著的延迟瓶颈。在本文中,我们提出了分块聚类注意力(BC注意力)来加速VGGT中的全局注意力层。通过将聚类限制在硬件友好的邻域块内,BC注意力减少了查询聚类的计算开销以及片上和片外存储器之间昂贵的数据移动。这使得BC注意力能够扩展到长序列,并在GPU上实现实际的延迟改进。此外,我们引入了一种哈希超平面校准方法和一种基于阈值的误差补偿方法,以高效减少聚类误差,这是当前聚类注意力机制中的一个瓶颈。总体而言,我们在GPU上的实验表明,校准的BC注意力将全局注意力层加速了2.10-2.63倍,整个骨干网络加速了1.77-2.35倍,且对于大场景性能损失可忽略不计(1%)。在较小的性能损失(<5%)下,校准的BC注意力进一步实现了全局注意力层2.26-2.87倍的延迟改进,以及骨干网络1.90-2.55倍的改进。

英文摘要

The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with extremely long sequences that causes a significant latency bottleneck. In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT. By limiting the clustering within HW-friendly neighborhood blocks, BC attention reduces the computation overhead of query clustering as well as the costly data movement between on- and off-chip memory. This enables BC attention to scale to long sequences and deliver practical latency improvements on GPUs. Moreover, we introduce a hashing hyperplane calibration method and a threshold-based error compensation method to reduce clustering errors efficiently, which is a bottleneck in the current clustered attention mechanism. Overall, our experiments on GPU demonstrate that calibrated BC attention accelerates the global attention layers by 2.10-2.63$\times$ and the whole backbone by 1.77-2.35$\times$ with negligible loss (1%) for large scenes. With a small performance loss (< 5%), calibrated BC attention further achieves a 2.26-2.87$\times$ latency improvement on the global attention layers and a 1.90-2.55$\times$ improvement on the backbone.

CommentsAccepted to IJCNN26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑