arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15260cs.CVcs.AI

VGGT-Align:为长序列三维重建连接局部重建与全局一致性

VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction

Wei Zhang, Yihang Wu, Songhua Li, Qi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对长序列三维重建的尺度漂移问题,提出VGGT-Align框架,通过SGIA和测试时适配策略,在多基准上实现顶尖性能,降低绝对轨迹误差最多32%

中文摘要 AI 辅助

保持全局几何一致性是长序列三维重建的核心挑战,其中尺度漂移是最关键的失效模式。在基于分块的推理流程中,连续Sim(3)对齐中的尺度自由度未受约束,导致估计误差呈倍数累积,扭曲全局轨迹与点云几何。我们提出一种尺度一致性增强框架,其核心见解是:在驾驶场景等结构化环境中,由环境规律性产生的几何量在各时间片段间固有不变,各分块对这些量的测量差异可直接暴露分块间的尺度漂移。我们提出场景几何不变锚定(SGIA),通过由粗到细的鲁棒估计从每个分块的预测点云中提取主导几何不变量,利用其跨分块一致性建立独立于点云配准的尺度约束,明确将7自由度(7-DoF)Sim(3)对齐退化为6自由度(6-DoF)刚体变换,从源头上切断链式尺度误差传播。我们进一步引入轻量测试时适配策略,仅通过多目标自监督微调归一化层参数,逐步改进序列内的分块预测。两个模块均为即插即用,无需离线重新训练。在多个长序列基准上的实验表明,该方法达到了顶尖性能,绝对轨迹误差降低最多32%,且在轨迹稳定性与重建质量上有显著提升。代码:this https URL

英文摘要

Maintaining global geometric consistency is a central challenge in long-sequence 3D reconstruction, with scale drift being the most critical failure mode. In chunk-based inference pipelines, the scale degree of freedom in sequential Sim(3) alignment is left unconstrained, causing estimation errors to compound multiplicatively and distort global trajectories and point cloud geometry. We present a scale-consistency enhancement framework built on a key insight: in structured environments such as driving scenes, geometric quantities arising from environmental regularity remain inherently invariant across temporal segments, and discrepancies in their per-chunk measurements directly expose inter-chunk scale drift. We propose Scene Geometric Invariant Anchoring (SGIA), which extracts dominant geometric invariants from each chunk's predicted point cloud via coarse-to-fine robust estimation and exploits their cross-chunk consistency to establish scale constraints independent of point cloud registration, explicitly degenerating 7-DoF Sim(3) alignment into 6-DoF rigid-body transformation and severing chain-wise scale error propagation at its source. We further introduce a lightweight test-time adaptation strategy that fine-tunes only normalization-layer parameters via multi-objective self-supervision, progressively improving intra-chunk predictions along the sequence. Both modules are plug-and-play and require no offline retraining. Experiments on multiple long-sequence benchmarks demonstrate state-of-the-art performance, reducing absolute trajectory error by up to 32% with significant gains in trajectory stability and reconstruction quality. Code: https://github.com/WZ-CS/VGGT-Align

发表机构

  • Northwestern Polytechnical University(西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑