arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboSeg:基于单眼在手相机的机器人操作的在线部件级语义重建

RoboSeg: Online Part-Level Semantic Reconstruction for Robotic Manipulation via a Single Eye-in-Hand Camera

Zhaochen Lan, Mengxiang Lin

arXiv 2608.09778首次发表:更新:

发表机构

School of Mechanical Engineering and Automation, Beihang University(北京航空航天大学机械工程及自动化学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RoboSeg是无需CAD模型的部件级语义重建系统,结合VLM、TSDF融合与SAM3,实现机器人操作的部件语义索引,在24次物理试验中21次完成综合任务。

AI 中文摘要

机器人操作需要感知系统能够识别可操作部件,如把手、轮辋、扳机和工具尖端,而非仅识别物体类别或点云。本文提出RoboSeg,这是一种部件级语义重建系统,它将视觉语言模型(VLM)的功能部件发现、异步在线RGB-D语义重建以及面向任务的抓取生成相结合,无需CAD模型或预扫描网格。RoboSeg在初始RGB观测上查询VLM以获取紧凑的功能部件提示,随后通过两个异步流进行扫描:高频几何线程用于RGB-D里程计和截断符号距离函数(TSDF)融合,以及关键帧触发的语义线程用于SAM3部件掩码。投影的掩码通过体素级时间投票融合为持久的部件标签点云;RoboSeg利用该地图将AnyGrasp六自由度候选分配给语义部件,并选择与任务相关部件标签一致的抓取。RoboSeg在手动标注物体上达到83.4%的平均部件交并比(mIoU);在涉及4个物体和8个任务的24次物理试点试验中,所选抓取在所有试验中都接触到请求的部件,并在24次试验中完成了21次综合任务。这些结果表明RoboSeg是面向任务的操作的语义索引层,AnyGrasp则保留为候选生成器。

英文摘要

Robotic manipulation requires perception systemsthat identify actionable parts such as handles, rims, triggers,and tool tips, not merely object categories or point clouds. This paper presents RoboSeg, a part-level semantic reconstructionsystem that links vision-language model (VLM) functional-partdiscovery, asynchronous online RGB-D semantic reconstruc-tion, and task-oriented grasp generation without requiring CAD models or pre-scanned meshes. RoboSeg queries a VLM onthe initial RGB observation to obtain compact functional part prompts, then scans with two asynchronous streams: a high-frequency geometry thread for RGB-D odometry and truncated signed distance function (TSDF) fusion, and a keyframe-triggered semantic thread for SAM3 part masks. Projectedmasks are fused by voxel-level temporal voting into a persistentpart-labeled point cloud; RoboSeg uses this map to assign AnyGrasp 6-DoF candidates to semantic parts and select grasps consistent with the task-relevant part label. RoboSeg reaches 83.4% mean part intersection-over-union (mIoU) over manually labeled objects; in a 24-trial physical pilot across fourobjects and eight tasks, the selected grasp contacts the requestedpart in all trials and achieves 21/24 combined task successes.These results characterize RoboSeg as a semantic indexing layerfor task-conditioned manipulation, with AnyGrasp retained asthe proposal generator.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑