AI 中文总结
QuacamFM提出四元数约束流匹配框架,通过球面线性插值保持单位四元数约束,提升稀疏视图相机位姿估计精度,优于扩散模型和经典SfM方法。
AI 中文摘要
从多视角图像进行相机位姿估计仍然是计算机视觉中的一个挑战。传统方法通常使用运动恢复结构(SfM)结合光束法平差来解决这一问题。然而,由于几何约束不足,从稀疏视图估计的相机位姿固有地存在模糊性。最近的工作利用概率模型(如扩散模型)生成多个相机位姿假设,从而更好地捕捉这种不确定性。这些方法大多使用单位四元数表示相机旋转,但在生成过程中将其视为无约束的4D向量,从而忽略了四元数的单位范数约束。无约束的四元数会产生不光滑且次优的生成轨迹。为此,我们提出了QuacamFM,一个用于相机位姿估计的四元数约束流匹配框架,在整个流轨迹中保持单位四元数表示。我们使用平滑的球面线性插值设计四元数流的最优传输。在CO3Dv2上的实验证明了我们的方法在相机位姿精度上优于基于扩散的方法和经典SfM方法。我们进一步表明,在稀疏视图相机位姿估计中,我们的四元数约束公式优于将标准流匹配直接应用于4D四元数向量的朴素方法。最后,观察到QuacamFM在跨数据集和野外示例中具有良好的泛化能力。
英文摘要
Camera pose estimation from multi-view images remains a challenge in computer vision. Traditional methods often address this problem using Structure-from-Motion (SfM) with bundle adjustment. However, camera poses estimated from sparse views are inherently ambiguous due to insufficient geometric constraints. Recent work leverages probabilistic models, such as diffusion models, to generate multiple camera pose hypotheses and therefore capture this uncertainty better. Most of these methods represent camera rotations using unit quaternions, but treat them as unconstrained 4D vectors during the generative processes, thereby ignoring the unit-norm constraint of quaternions. Unconstrained quaternions create non-smooth and suboptimal generation trajectories. To this end, we propose *QuacamFM*, a quaternion-constrained flow matching framework for camera pose estimation that preserves unit quaternion representations throughout the entire flow trajectory. We design the optimal transport of the quaternion flows using smooth spherical linear interpolation. Experiments on CO3Dv2 demonstrate our method's advantage in camera pose accuracy over diffusion-based methods and classical SfM approaches. We further show that our quaternion-constrained formulation outperforms the naive application of standard flow matching to 4D quaternion vectors on sparse-view camera pose estimation. Finally, it is observed that QuacamFM generalizes well across datasets and in-the-wild examples.