基于分级非凸性的无三角化光束平差法,用于从粗先验中优化相机位姿
Triangulation-Free Bundle Adjustment with Graduated Non-Convexity for Camera Pose Refinement from Coarse Priors
浏览论文内容
中文总结 AI 辅助
本文提出无三角化的分级非凸性光束平差法,避免先验误差固化,在单CPU上实现高精度相机位姿优化,抗扰动能力优于传统方法,效率远高于学习型优化器。
中文摘要 AI 辅助
移动增强现实(AR)框架会为每一次普通手机拍摄附加一个度量位姿先验,在CPU上将其低成本转换为重建级位姿,是新视角合成前的关键步骤。优化器至少应保证不会让准确的先验恶化,但现有主流优化器却未能做到:在15个ScanNet++ iPhone房间场景采集数据上,COLMAP三角化加先验初始化的光束平差法,使所有15个ARKit准确先验的场景平均误差从0.55度升至0.74度,原因在于初始化阶段的三角化——结构从先验中三角化后再进行优化,导致先验误差被固化到优化器所依赖的结构中。本文移除了三角化步骤:每个关键点在其自身反投影射线上拥有一个标量深度,每个匹配贡献两个对称交叉投影残差,因此结构在每次迭代中都会被重新表达。该方法在330次带扰动的房间测试中,将房间先验误差保持在0.57度且无失败案例;在物体尺度上,从0.456度的先验出发,达到0.265度/1.80毫米的精度,单CPU上每个场景中位数耗时10秒,而学习型优化器需2.5 GPU小时。由于未固化结构,目标函数支持分级非凸性,可衡量缺陷程度:传统优化器在先验误差超过1-2度(仅略高于真实ARKit先验)后会失效,无传统优化器能承受32度误差;本文方法在16度/80毫米的425次测试中全部恢复,在32度/160毫米的测试中恢复85%,可承受带扰动房间的32度误差。标称物体尺度精度相当但未更优,在自身噪声底的基准上,传统光束平差法是文献中缺失的强基线;零扰动下每个求解器已有1个场景失败案例,基于位置先验的重映射方法虽匹配先验帧但会丢弃先验,无法利用值得保留的先验或进行热启动。
英文摘要
Mobile AR frameworks attach a metric pose prior to every casual phone capture, and turning it into reconstruction-grade poses cheaply on CPU is the step before novel-view synthesis. The least a refiner owes an accurate prior is not to make it worse. The workhorse refiner does. On 15 ScanNet++ iPhone room captures, COLMAP triangulation plus prior-seeded bundle adjustment degrades an accurate ARKit prior in all 15, 0.55 degrees to 0.74 degrees by scene-mean. The cause is the seeding. Structure is triangulated from the prior before anything is optimized, so the prior's error is baked into the structure the optimizer trusts. We remove the triangulation. Every keypoint owns a scalar depth along its own back-projected ray and each match contributes two symmetric cross-projection residuals, so structure is re-expressed at every iterate. The same solve holds the room prior at 0.57 degrees and never fails in 330 perturbed room runs, and at object scale reaches 0.265 degrees/1.80 mm from a prior at 0.456 degrees in a median of 10 s per scene on one CPU, against 2.5 GPU-hours for a learned refiner. Because no structure is committed, the objective also admits graduated non-convexity, which measures how deep the defect goes. Classical refinement collapses past 1-2 degrees of prior error, barely beyond a real ARKit prior, and no classical refinement arm survives 32 degrees. Ours recovers 425 of 425 runs through 16 degrees/80 mm and 85% at 32 degrees/160 mm, and perturbed rooms through 32 degrees. Nominal object-scale accuracy is on par rather than better, on a benchmark at its own noise floor, where classical bundle adjustment is a strong baseline absent from the literature. One scene fails for every solver already at zero perturbation. Re-mapping from position priors matches us in the prior's frame but discards it, so it cannot exploit a prior worth keeping or be warm-started.