arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoMo:机器人操作中基于时空动作标记化的拨号运动模式

MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization

Yuhan Hu, Hugues Thomas, Peide Huang, Mouli Sivapurapu, Benoit Landry, Arto Kivila

arXiv 2607.26315首次发表:更新:

AI 中文总结

MoMo是两阶段模仿学习框架,可学习跨任务复用的运动模式,能生成不同模式的操作行为并迁移未见过的模式,保持任务成功率,实现组合泛化。

AI 中文摘要

为在多样场景中高效运行,机器人不仅需精准完成操作任务,还需根据任务、对象及交互场景调整动作执行方式。本文探究这种执行层面的差异能否作为跨任务共享的可复用行为因子被学习。我们提出MoMo,这是一个两阶段模仿学习框架,由时空动作标记器和行为克隆Transformer构成,输入为任务与连续运动模式条件。在6项真实机器人操作任务中,调整该条件可生成平稳、动态及中间行为,人类评估者可区分这些行为,且它们在关节速度、加速度及末端执行器接近俯仰角上存在差异。对于仅以一种模式演示的任务,MoMo可迁移未见过的请求模式,同时基本保持任务成功率。综上,这些结果为对未见过的任务-模式组合的组合泛化提供了证据,表明运动模式可跨任务复用,以控制操作技能的执行方式。

英文摘要

To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present \textbf{MoMo}, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode condition as inputs. Across six real-robot manipulation tasks, varying this condition produces steady, dynamic, and intermediate behaviors that human raters can distinguish and that differ in joint speed, acceleration, and end-effector approach pitch. On tasks demonstrated in only one mode, MoMo transfers the unseen requested mode while largely preserving task success. Together, these results provide evidence of compositional generalization to unseen task--mode combinations and show that motion mode can be reused across tasks to control how a manipulation skill is performed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑