发表机构
Karlsruhe Institute of Technology; University of Bonn; Technical University of Darmstadt; Lamarr Institute for Machine Learning and Artificial Intelligence; Center for Robotics, Bonn(卡尔斯鲁厄理工学院; 波恩大学; 达姆施塔特工业大学; 拉马尔机器学习与人工智能研究所; 波恩机器人中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对杂乱狭窄空间的机器人建图难题,提出MS-MEM框架,整合主动视点选择、推搡与抓取,引入附带干扰约束,实现更高建图精度并减少场景干扰。
AI 中文摘要
在货架等狭窄、杂乱空间中实现准确的场景理解对服务机器人至关重要,因为许多日常任务要求它们可靠地定位和取回物体。然而,由于严重的遮挡、受限的可达性以及需要避免场景发生过多变化,这一任务仍具挑战性。本文提出多技能操纵增强建图(Multi-Skill Manipulation-Enhanced Mapping,MS-MEM),这是一种用于感知不确定性感知建图的证据框架,整合了主动视点选择、物体推搡与抓取操作。MS-MEM将场景级度量-语义证据信念估计器与不确定性感知抓取表示相结合,该表示通过一种新型全证据抓取估计器学习,可同时对抓取可行性和方向不确定性进行建模。在我们的框架中,候选感知与操纵操作通过统一的动作选择流程,使用共同的信息增益准则进行评估;对于操纵操作,我们进一步引入了附带干扰约束(collateral disturbance constraint,CDC),以避免对场景信念的置信区域造成过多改变。这使MS-MEM能够选择有效降低地图不确定性同时限制附带场景变化的操作。实验结果表明,与忽略场景干扰的单技能基线和无约束基线相比,MS-MEM在实现更高建图精度的同时大幅减少了场景干扰,凸显了主动视点选择、推搡与抓取操作的协同效应。
英文摘要
Accurate scene understanding in confined, cluttered spaces such as shelves is essential for service robots, as many everyday tasks require them to locate and retrieve objects reliably. Yet, it remains challenging due to severe occlusions, restricted accessibility, and the need to avoid excessive scene changes. In this paper, we propose Multi-Skill Manipulation-Enhanced Mapping (MS-MEM), an evidential framework for uncertainty-aware mapping that integrates active viewpoint selection, object pushing, and grasping. MS-MEM combines scene-level metric-semantic evidential belief estimators with an uncertainty-aware grasp representation. This representation is learned using a novel full-evidential grasp estimator that models both grasp affordance and orientation uncertainty. In our framework, candidate perception and manipulation actions are evaluated within a unified action selection pipeline using a common information gain criterion. For manipulation actions, we further introduce a collateral disturbance constraint (CDC) that discourages excessive changes to confident regions of the scene belief. This enables MS-MEM to select actions that effectively reduce map uncertainty while limiting collateral scene changes. Experimental results show that, compared with single-skill and unconstrained baselines that ignore scene disturbance, MS-MEM achieves higher mapping accuracy while substantially reducing scene disturbance, highlighting the synergistic effects of active viewpoint selection, push, and grasp actions.
Commentsunder review