arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

本体感觉草图作为生成式动作策略的长时程意图

Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies

Fangyuan Wang, Songhao Huang, Haoxiang Sun, Shipeng Lyu, Chengyang He, Anqing Duan, Peng Zhou, David Navarro-Alarcon

arXiv 2610.02759首次发表:更新:

发表机构

Hong Kong Polytechnic University; National University of Singapore; Mohamed bin Zayed University of Artificial Intelligence; Great Bay University(香港理工大学; 新加坡国立大学; 穆罕默德·本·扎耶德人工智能大学; 大湾区大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出本体感觉动作模型(PAM),通过生成紧凑的关节空间草图作为长时程意图,在Transformer去噪器中联合生成动作,提升机器人任务成功率。

AI 中文摘要

生成式机器人策略能够预测短时动作片段,但缺乏明确的长时程意图。近期方法通过语言计划、子目标图像或视频预测来展现更长时程的结构,但这些方法生成成本高昂,且仍需转化为机器人运动。直接预测未来机器人运动可避免这种转化,但密集的、按时间索引的轨迹需要大量参数以覆盖整个剩余任务,且在短时程内,它主要重复动作片段,对动作生成的指导作用有限。我们提出了本体感觉动作模型(PAM),该模型在单个Transformer去噪器中联合生成机器人剩余关节空间路径的紧凑、无时间戳草图,以及密集的可执行动作片段。草图以弧长而非时间进行参数化,捕捉不受执行时间影响的几何意图。块因果注意力和交错去噪调度保持了从草图到动作的定向依赖,确保动作令牌在采样过程中逐步基于更清晰的草图进行条件生成。在仿真中,PAM在Push-T和LIBERO-Long任务上优于仅基于动作的对应模型;在四个真实世界的双臂任务中,其成功率从47.5%提升至75.0%。项目页面:此https链接

英文摘要

Generative robot policies predict short action chunks but lack explicit long-horizon intent. Recent methods expose longer-horizon structure through language plans, subgoal images, or video forecasts, which are costly to generate and still need to be translated into robot motion. Predicting future robot motions avoids this translation, but a dense, time-indexed trajectory requires numerous parameters to cover the full remaining task, and over a short horizon it largely repeats the action chunk and adds little guidance for action generation. We propose Proprioceptive Action Models (PAM), which jointly generate a compact, timing-free sketch of the robot's remaining joint-space path and a dense executable action chunk within a single transformer denoiser. The sketch parameterizes the path by arc length rather than time, capturing geometric intent invariant to execution timing. Block-causal attention and a staggered denoising schedule maintain directed sketch-to-action dependence, ensuring the action tokens condition on a progressively cleaner sketch throughout sampling. In simulation, PAM improves over its action-only counterparts on Push-T and LIBERO-Long; on four real-world bimanual tasks, it raises success from 47.5% to 75.0%. Project page: https://nicehiro.github.io/pam_dp/

Comments19 pages. Project page: https://nicehiro.github.io/pam_dp/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑