arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10534cs.RO

RoboPrompt:基于稀疏人类输入的直观机器人策略引导

RoboPrompt: Intuitive Robot Policy Steering with Sparse Human Input

Yanwen Zou, Chenyang Shi, Guoxuan Xu, Wenye Yu, Wendi Chen, Ye Pan, Cewu Lu, Chuan Wen

首次发表
浏览论文内容

中文总结 AI 辅助

提出RoboPrompt,一种无需修改策略架构的轻量级引导系统,通过稀疏人类输入控制动作生成,经DAgger迭代后显著提升多种策略的成功率并减少人类干预。

中文摘要 AI 辅助

通过模仿学习训练得到的端到端机器人策略仍受限于数据多样性的不足,这使得其在真实世界场景中的可靠零样本部署面临挑战。共享自主方法通过遥操作实现人类纠正,但专用硬件和操作员培训阻碍了其大规模部署。其他方法将人类指导作为额外的策略输入,通常需要对架构进行修改并进行专门的训练以实现可引导性,这限制了它们在不同策略间的适用性。我们提出了RoboPrompt,一个通用、轻量级的机器人策略引导系统,使用户能够通过直观的稀疏输入(包括手绘轨迹、目标点和粗略的方向指令)来引导策略行为。RoboPrompt将人类意图翻译与底层策略解耦:一个可复用模块将人类指导转换为动作草稿,这些草稿再通过基础策略的扩散或流匹配动力学进行细化。通过在噪声空间中控制动作生成,RoboPrompt在人类意图与策略先验之间取得平衡,而无需修改基础策略架构或对其进行可引导性微调。实验表明,该系统在Diffusion Policy、$\pi_{0.5}$和FastWAM上均实现了有效的引导。我们进一步利用引导的轨迹通过DAgger进行在线策略改进。经过2-3轮迭代,$\pi_{0.5}$在三个任务上的平均成功率提升了15.5%,在Insert Bread任务上三种策略(Diffusion Policy、$\pi_{0.5}$、FastWAM)的平均成功率提升了21.3%,而平均人类干预次数分别减少了44.0%(从2.86降至1.60)和81.9%(从2.60降至0.47)。

英文摘要

End-to-end robot policies trained through imitation learning remain constrained by limited data diversity, making reliable zero-shot deployment in real-world settings challenging. Shared-autonomy methods enable human correction through teleoperation, but specialized hardware and operator training hinder deployment at scale. Other approaches incorporate human guidance as additional policy inputs, often requiring architectural changes and dedicated training for steerability, which limits their applicability across policies. We present RoboPrompt, a general-purpose, lightweight robot policy steering system that enables users to guide policy behavior through intuitive, sparse inputs, including drawn traces, target points, and coarse directional instructions. RoboPrompt decouples human-intention translation from the underlying policy: a reusable module converts human guidance into action drafts, which are refined through the diffusion or flow-matching dynamics of the base policy. By controlling action generation in noise space, RoboPrompt balances human intent with the policy prior without modifying the base policy architecture or fine-tuning it for steerability. Experiments demonstrate effective steering across Diffusion Policy, $π_{0.5}$, and FastWAM. We further use steered rollouts for online policy improvement through DAgger. After 2-3 rounds of iteration, average success rates increase by 15.5\% for $π_{0.5}$ across three tasks and by 21.3\% across three policies(Diffusion Policy, $π_{0.5}$, FastWAM) on the Insert Bread task, while average human intervention counts decrease by 44.0\% (2.86 to 1.60) and 81.9\% (2.60 to 0.47), respectively.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Noematrix Ltd.(Noematrix有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑