DreamSat-Pose:基于单视图3D重建和学习的2D-3D特征匹配的航天器姿态估计
DreamSat-Pose: Spacecraft Pose Estimation from Single-View 3D Reconstructions and Learned 2D-3D Feature Matching
- Wellesley College(韦尔斯利学院)
- Massachusetts Institute of Technology(麻省理工学院)
- Politecnico di Milano(米兰理工大学)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Universidad Politécnica de Madrid(马德里理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对未知航天器对象,本文提出新框架,先利用冻结的DINOv3提取图像特征、动态图卷积神经网络编码器计算几何特征,再经双流变压器匹配器细化描述符,通过学习2D-3D对应关系估计姿态,在数据集上评估效果良好,泛化能力强。
AI中文摘要:
6自由度姿态估计在自主交会和接近操作中是一项关键任务。在目标未知的情况下,由于要与目标形状模型重建相结合,该任务具有挑战性。本文提出了一种用于未知航天器对象的单次形状和姿态估计的新框架。给定单张图像,先重建目标的3D形状模型,然后通过学习密集的2D-3D对应关系估计相对六自由度姿态。利用冻结的DINOv3视觉变压器提取图像特征,通过可训练的动态图卷积神经网络编码器从重建点云计算几何特征。双流变压器匹配器通过交替自注意力和交叉注意力细化描述符,产生软对应关系并传递给透视n点求解器进行姿态恢复。在SPE3R数据集上评估该方法,并将FoundationPose作为当前最先进能力的代表性基线。结果表明,仅使用单张图像和重建几何就能实现可靠的姿态估计,平均指向误差为0.157度,对未见航天器具有很强的泛化能力。
英文摘要:
6-DoF pose estimation is a critical task in autonomous rendezvous and proximity operations. In the case of an unknown target, this task becomes challenging as it shall be paired with the reconstruction of the target shape model. In this article, we propose a novel framework for single-shot shape and pose estimation of unknown spacecraft objects. Given a single image, we first reconstruct a 3D shape model of the target, then estimate the relative six-degrees-of-freedom pose by learning dense 2D-3D correspondences. The image features are extracted using a frozen DINOv3 vision transformer, while the geometric features are computed from the reconstructed point cloud using a trainable dynamic graph convolutional neural network encoder. A dual-stream transformer matcher refines descriptors through alternating self- and cross-attention, producing soft correspondences that are passed to a Perspective-$n$-Point solver for pose recovery. We evaluate the method on the SPE3R dataset and consider FoundationPose as a representative baseline for current state-of-the-art capabilities. Results show reliable pose estimates achieving 0.157 degrees mean pointing error using only a single image and reconstructed geometry, demonstrating strong generalization to unseen spacecraft.