发表机构
University of British Columbia(不列颠哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VGGT全局注意力中的架构冗余,提出计算自适应混合头模型VGGT-Prime,通过轻量路由器动态分配计算模式,实现8倍加速并保持重建质量,且可与令牌合并互补达14倍加速。
AI 中文摘要
前馈视觉几何模型,如视觉几何接地变换器(VGGT),近期已能够从多视图图像直接进行三维重建。尽管这些模型性能优异,但由于其全局注意力机制,计算复杂度随输入视图数量呈二次方增长,导致长序列输入时延迟显著。近期有一些加速VGGT的尝试,但主要集中于通过令牌合并或键/值稀疏化来减少令牌冗余。我们的工作从不同角度解决这一瓶颈,研究视觉几何变换器中的架构冗余。我们发现VGGT全局注意力层中的多头注意力模块存在显著的架构冗余,仅有一部分头携带关键的几何信息。基于此观察,我们提出VGGT-Prime,一种计算自适应的混合头模型,通过解决该冗余来加速视觉几何变换器,同时保持有竞争力的重建质量。VGGT-Prime的关键思想是使用轻量级路由器估计每个全局注意力头的适当计算水平,然后动态地将每个头分配到不同的计算模式。在多个数据集上的大量实验表明,VGGT-Prime相比VGGT可实现8倍推理加速,同时在相机位姿、深度和点云预测上保持有竞争力的性能。我们进一步展示VGGT-Prime与现有加速方法(如令牌合并)互补,相比VGGT可将推理速度进一步提升至14倍。我们工作的概述可在项目页面获取。
英文摘要
Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-view images. Despite their promising performance, these models scale quadratically with the number of input views due to their global attention mechanism, resulting in substantial latency for long sequence inputs. There have been some recent efforts to accelerate VGGT, but they primarily focus on reducing \emph{token redundancy} through token merging or key/value sparsification. Our work resolves this bottleneck from a different perspective by investigating \emph{architectural redundancy} in visual geometry transformers. We show that the multi-head attention modules in VGGT's global-attention layers contain substantial architectural redundancy, with only a subset of heads carrying critical geometric information. In light of this observation, we propose VGGT-Prime, a compute-adaptive mixture-of-heads model that resolves this redundancy to accelerate visual geometry transformers while maintaining competitive reconstruction quality. The key idea of VGGT-Prime is to estimate the appropriate computation level for each global-attention head using a lightweight router and then dynamically assign each head to different computation modes. Extensive experiments on multiple datasets demonstrate that VGGT-Prime can achieve an {$8\times$} inference speedup over VGGT while maintaining competitive performance on camera pose, depth, and point-cloud predictions. We further show that VGGT-Prime is complementary to existing acceleration methods, such as token merging, further improving inference speed by up to $14{\times}$ over VGGT. An overview of our work is available on our \href{https://vggt-prime.github.io}{project page}.
CommentsTechnical Report