发表机构
Korea Advanced Institute of Science and Technology(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出VGGT-Bridge,通过添加粗步长跳跃边连接非相邻分块,并利用反向粗块校正漂移,在不重训练下显著降低长序列重建的ATE误差。
AI 中文摘要
前馈式视觉几何变换器(如VGGT)能在单次前向传播中从图像重建稠密三维结构,简化了多视图三维重建。然而,其二次方注意力复杂度使其难以扩展到包含数千帧的长序列。分块对齐框架通过将长序列分割为重叠块并将局部重建拼接成位姿图来解决此问题。但现有方法仅连接顺序相邻的块,导致逐帧小误差沿链累积成大尺度漂移。为超越顺序边,我们提出VGGT-Bridge,在不重新训练的情况下添加长程跳跃边,直接约束非相邻块。通过在稀疏采样的粗块上运行VGGT,每个粗块将远距离细块桥接为单一直接约束。我们进一步将VGGT的首帧尺度偏差转化为漂移校正,通过反向馈送选定的粗块,且循环感知策略使该反向操作与现有闭环兼容。VGGT-Bridge在KITTI Odometry上相比SwiftVGGT基线将ATE降低28.3%,在Virtual KITTI上降低18.8%,在Waymo Open上降低10.0%,在所有分块对齐方法中达到最佳性能。
英文摘要
Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences with thousands of frames. Chunk-and-align frameworks address this by splitting a long sequence into overlapping chunks and stitching their local reconstructions into a pose graph. Yet existing methods connect only sequentially adjacent chunks, so small per-frame errors accumulate along the chain into large-scale drift. To move beyond sequential edges, we propose VGGT-Bridge, which adds long-range skip edges that directly constrain non-adjacent chunks without retraining. By running VGGT on sparsely sampled coarse chunks, each coarse chunk bridges distant fine chunks into a single direct constraint. We further turn VGGT's first-frame scale bias into a drift correction by feeding selected coarse chunks in reverse, and a loop-aware policy keeps this reversal compatible with existing loop closures. VGGT-Bridge reduces ATE by 28.3% on KITTI Odometry, 18.8% on Virtual KITTI, and 10.0% on Waymo Open over the SwiftVGGT baseline, achieving the best performance among all chunk-and-align methods.
CommentsAccepted to ACCV 2026