arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双臂移动操作任务的自动标注

Automatic Labelling for Bimanual Mobile Manipulation

Yupu Lu, Jia Pan

arXiv 2609.24059首次发表:更新:

发表机构

The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种结合轨迹分割与视觉-语言推理的自动标注流水线,用于双臂移动操作任务,在29个真实任务上验证了其生成结构化标注的有效性。

AI 中文摘要

具有语义意义的子任务标签可以为长时程策略提供有用的上下文,但自动识别可靠的时间边界和广泛的语义描述用于标注仍然困难。我们提出了一种自动标注流水线,将时间定位分配给确定性轨迹分析,将语义解释分配给视觉-语言(VL)推理。该流水线将同步的运动学信号分割成阶段,执行阶段定位的VL推理以描述内容,并聚合基础、左臂和右臂动作的输出。我们主要在29个真实的Galaxea双臂移动操作任务上评估该流水线。重复VL推理三次,在选定任务上首次产生相同输出值的比例为87.4%。随后,九名参与者对所有29个任务的标注阶段进行审查,对时间划分(90.5%)、身体标签(90.7%)和手臂标签(78.7%)的接受度均为正向。结果表明,分割-VL设计可以生成结构化标注,同时保留异步双臂行为,为更丰富的语义子任务识别和基于状态的验证提供了基础。

英文摘要

Semantically meaningful subtask labels can provide useful contexts for long-horizon policies, but automatically identifying both reliable temporal boundaries and broad semantic descriptions for annotations remains difficult. We present an automatic labelling pipeline that assigns temporal localisation to deterministic trajectory analysis and semantic interpretation to vision-language (VL) reasoning. The pipeline segments synchronised kinematic signals into phases, performs phase-localised VL reasoning to describe the contents, and aggregates the outputs for the base, left arm, and right arm actions. We evaluate this pipeline primarily on 29 real Galaxea bimanual mobile-manipulation tasks. Repeating the VL reasoning three times first produces the same output value for 87.4% on selected tasks. A review by nine participants across all 29 tasks then judgements on the labelled phases and shows positive acceptance of temporal divisions (90.5%), body labels (90.7%), and arm labels (78.7%). The results indicate that the segmentation-VL design can produce structured annotations while preserving asynchronous bimanual behaviour, providing a basis for richer semantic subtask identification and state-based verification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑