arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VGGT 家族是否需要其所有层?

Does the VGGT Family Need All Its Layers?

Fengyi Zhang, Holger Caesar, Xiangyu Sun, Zheng Zhang, Zi Huang, Yadan Luo

arXiv 2609.36842首次发表:更新:

发表机构

The University of Queensland; Delft University of Technology; Harbin Institute of Technology(昆士兰大学; 代尔夫特理工大学; 哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究 VGGT 等前馈几何模型的层冗余,发现可移除层集中于早期和后期区域,提出基于 CKA 的剪枝代理和闭式线性校准,在减少 44% 聚合器参数的同时保持准确性。

AI 中文摘要

前馈几何模型的哪些层对于保持相机姿态和稠密三维结构是必需的?我们研究了 VGGT、π³ 和 VGGT-Ω 中的层冗余:共 3018 个剪枝配置,在四个室内和室外数据集上,针对七个相机姿态和稠密几何指标进行了评分。我们得出以下四个发现:(i)可移除的层聚集在两个冗余区域:一个占主导的早期区域和一个较窄的后期区域,而跨越中间层的删除始终更具破坏性。这种重复模式在模型、数据集和指标中均成立,与文献中通常报道的中后期冗余形成对比。(ii)在这些区域内,我们观察到删除两个区间的联合退化大约等于它们各自退化的总和,这将剪枝搜索的模型评估次数从 O(L^4) 减少到 O(L^2),其中 L 是聚合器深度。(iii)我们发现 CKA 为区间退化提供了更便宜的基于表示的代理,在剪枝质量和校准成本之间提供了实用的权衡。(iv)闭式线性校准在剪枝后无需端到端重训练即可恢复准确性。最小二乘分析表明,对特殊 token 和补丁 token 使用共享映射通常会导致额外的重建损失,这促使我们采用 token 感知的恢复。仅在 100 个校准场景上拟合的恢复映射可以泛化到保留场景和未见数据集。由此产生的模型将聚合器参数减少高达 44%,同时保持与完整模型相当的准确性。代码和实验结果将在我们的项目页面提供:此 https URL

英文摘要

Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, $π^3$, and VGGT-$Ω$: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Removable layers cluster in two redundancy regions: a dominant early region and a narrower late one, while deletions spanning the intervening layers are consistently more disruptive. This recurring pattern holds across models, datasets, and metrics, and contrasts with the middle-to-late redundancy commonly reported in the literature. (ii) Within these regions, we observe that the joint degradation from deleting two intervals is approximately the sum of their individual degradations, reducing the number of model evaluations for pruning search from $O(L^4)$ to $O(L^2)$, where $L$ is the aggregator depth. (iii) We find that CKA provides a cheaper representation-based proxy for interval degradation, offering a practical trade-off between pruning quality and calibration cost. (iv) Closed-form linear calibration recovers accuracy after pruning without end-to-end retraining. A least-squares analysis shows that using a shared map for special and patch tokens generally incurs excess reconstruction loss, motivating token-aware recovery. Recovery maps fitted on just 100 calibration scenes generalize to held-out scenes and unseen datasets. The resulting models reduce aggregator parameters by up to 44% while maintaining accuracy comparable to their intact counterparts. Code and experimental results will be available at our project page: https://xian-bei.github.io/vggt-family-layer-redundancy/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑