基于学习排序与验证的智能体相对相机位姿估计
Agentic Relative Camera Pose Estimation via Learned Ranking and Verification
浏览论文内容
中文总结 AI 辅助
PoseAgent通过可学习的排序与验证智能体动态编排多个位姿估计器,在多个基准上显著提升相对相机位姿估计精度。
中文摘要 AI 辅助
针对相机位姿估计,已有多种方法被开发出来,包括基于对应关系的方法、端到端位姿回归以及最新的三维几何基础模型。我们的关键观察是,对于宽基线、缺乏纹理、外观变化和遮挡等多样化的挑战,没有任何单一的估计器是最优的。进一步的分析揭示了在基准测试和单个图像对之间性能存在显著差异,不同的估计器表现出互补的优势。我们提出了PoseAgent,一个用于相对相机位姿估计的智能体框架,通过可学习的排序和验证动态地编排位姿估计器。给定一个图像对,一个特征分析智能体首先提取与位姿估计相关的外观、语义和几何特征,例如场景类型。然后,一个学习排序智能体根据图像对特征预测多个位姿估计器的相对能力。执行排名最高的估计器,其预测的位姿由一个学习验证智能体评估,该智能体估计相应的位姿误差。当验证失败时,PoseAgent自适应地调用排名较低的估计器,直到候选被接受或达到执行预算。对于位姿验证,我们的验证网络比之前的模型更准确地预测位姿误差。对于位姿估计,PoseAgent在ARKitScenes、MegaDepth、ScanNet++和RealEstate10K上,相比最强的独立估计器,AUC@5度指标分别提升了高达4.2%。在ARKitScenes上,PoseAgent也优于基于VLM的智能体,这些智能体包括一个具有相同验证器和回退策略的VLM排序器。这些结果证明了我们学习排序和验证的有效性。
英文摘要
A wide range of approaches have been developed for camera pose estimation, including correspondence-based methods, end-to-end pose regression, and recent 3D geometric foundation models. Our key observation is that no single estimator is optimal for diverse challenges, such as wide baselines, lack of texture, appearance changes, and occlusions. Further analysis reveals substantial performance variation across both benchmarks and individual image pairs, with different estimators exhibiting complementary strengths. We introduce PoseAgent, an agentic framework for relative camera pose estimation that dynamically orchestrates pose estimators through learnable ranking and verification. Given an image pair, a profiling agent first extracts appearance, semantic, and geometric features relevant to pose estimation, e.g., scene type. A learned ranking agent then predicts the relative competence of multiple pose estimators given the image-pair profile. The top-ranked estimator is executed, and its predicted pose is assessed by a learned verification agent that estimates the corresponding pose error. When verification fails, PoseAgent adaptively invokes lower-ranked estimators until a candidate is accepted or the execution budget is reached. For pose verification, our verification network predicts pose errors more accurately than prior models. For pose estimation, PoseAgent improves AUC@5 degree up to 4.2% over the strongest standalone estimator on each of ARKitScenes, MegaDepth, ScanNet++, and RealEstate10K. On ARKitScenes, PoseAgent also outperforms VLM-based agents, which include a VLM ranker with the same verifier and fallback policy. These results demonstrate the effectiveness of our learned ranking and verification.
发表机构
- Arizona State University(亚利桑那州立大学)
- University of California, Merced(加利福尼亚大学默塞德分校)
- Lund University(隆德大学)
- University of Nevada, Reno(内华达大学里诺分校)
机构由 AI 辅助整理,请以论文原文为准。