AI 中文总结
提出LoG-VGGT框架,通过跨窗口注意力与全局相机一致性精炼模块,在内存受限下实现长序列三维重建,提升深度精度与位姿稳定性。
AI 中文摘要
我们提出了LoG-VGGT,一种用于长序列三维重建的内存高效框架,它在局部时间建模与全局相机一致性之间取得平衡。我们的方法不依赖完全的全局注意力,而是在一小部分Transformer块中引入跨窗口注意力,从而在保持内存使用受限的同时,实现跨相邻时间窗口的有效信息传播。为缓解长期位姿漂移,我们进一步设计了一个全局相机一致性精炼模块,在该模块中,相机令牌通过交叉注意力与紧凑的寄存器令牌交互,以在整个序列中强制执行场景级约束。这一设计使得相机表示的联合优化成为可能,并显著提高了长时程位姿稳定性,而无需承担序列级注意力的高昂成本。大量实验表明,LoG-VGGT在多个长序列基准上实现了更高的深度精度和稳健的相机位姿估计,同时提供了具有竞争力的流式重建性能。
英文摘要
We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effective information propagation across adjacent temporal windows while keeping memory usage bounded. To mitigate long-term pose drift, we further design a global camera consistency refinement module, where camera tokens interact with compact register tokens via cross-attention to enforce scene-level constraints across the entire sequence. This design enables joint optimization of camera representations and significantly improves long-horizon pose stability without incurring the high cost of sequence-wide attention. Extensive experiments demonstrate that LoG-VGGT achieves improved depth accuracy and robust camera pose estimation across multiple long-sequence benchmarks, while delivering competitive streaming reconstruction performance.
Comments9 pages,4 figures