发表机构
University of Stuttgart(斯图加特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
QueryArt模型从单张RGB图像、2D查询点和相机内参估计关节参数,以查询点深度为单位回归3D几何,在多个基准上优于现有方法,并在真实机器人操纵试验中达到70.2%成功率。
AI 中文摘要
使机器人能够估计关节物体的运动学参数,将为其交互和操作解锁广泛的能力。这种估计必须基于机器人当前观察到的信息进行,通常仅是一张它从未见过的物体的单个RGB图像。当前的单图像方法将关节部分分割与关节估计耦合在一起,使得其预测容易受到漏检和错误部分关联的影响,并且它们回归度量3D几何,而单一视角只能确定到尺度。我们提出QueryArt,一种从单个RGB图像、一个2D查询点和相机内参估计关节参数的模型。QueryArt被训练为相对于查询点并以查询点深度为单位估计3D关节几何,这使得其目标仅从图像即可识别。查询点处的单个深度测量随后提供尺度并恢复度量参数。我们在一个精心混合的合成和真实世界关节数据集上训练QueryArt。我们在多个基准上评估QueryArt,并与最近的基线进行比较。QueryArt在大多数关节指标上优于最近的基线,包括在分布外数据上。为了展示模型在真实世界环境中的能力,我们在移动操纵器上评估QueryArt,跨越57次操纵试验,涵盖16个物体部分和五个视角类别,实现了70.2%的成功率。我们在以下网址提供代码和视频:此https URL。
英文摘要
Enabling robots to estimate the kinematic parameters of articulated objects unlocks a wide range of capabilities for interaction and manipulation. The estimation has to happen from the information the robot currently observes, often just a single RGB image of an object it has never seen before. Current single-image approaches couple articulation part segmentation with articulation estimation, making their predictions vulnerable to missed detections and incorrect part associations, and they regress metric 3D geometry that a single view fixes only up to scale. We present QueryArt, a model that estimates articulation parameters from a single RGB image, a 2D query point, and camera intrinsics. QueryArt is trained to estimate the 3D articulation geometry relative to the queried point and in units of its depth, which keeps its target identifiable from the image alone. A single depth measurement at the query point then supplies the scale and recovers the metric parameters. We train QueryArt on a curated mixture of synthetic and real-world articulation datasets. We evaluate QueryArt on several benchmarks and compare it against recent baselines. QueryArt outperforms recent baselines on most articulation metrics, including on out-of-distribution data. To demonstrate the model's capabilities in real-world settings, we evaluate QueryArt on a mobile manipulator across 57 manipulation trials spanning 16 object parts and five viewpoint classes, achieving a 70.2% success rate. We provide code and videos at: https://abwerby.github.io/queryart/
CommentsCode and video are available at https://abwerby.github.io/queryart/