arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越模仿:基于离线Q规划的机器人策略自改进

Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

Varun Giridhar, Anant Khandelwal, Jeremy A. Collins, Ignat Georgiev, Animesh Garg

arXiv 2608.21204首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Q规划方法,为大型视觉运动BC策略配备离线Q函数,通过仅微调Q函数实现自改进,在仿真和真实机器人任务中均显著提升机器人操作性能,且优于多种对比方法。

AI 中文摘要

行为克隆(BC)推动了机器人操作领域的显著进展,但它存在根本性局限:无法自改进——失败的策略若没有额外人类演示,就无法从失败中学习。强化学习微调提供了自改进的路径,但已证明难以扩展到支撑现代机器人策略的数十亿参数模型。我们提出Q规划,它为大型视觉运动BC策略配备了小型离线Q函数。由于Q函数估计价值而非模仿动作,它可以与BC策略在相同的成功演示上训练,之后还能吸收部署时的成功和失败轨迹,这是BC不具备的不对称性。我们利用这种不对称性,在推理时实现价值引导的动作选择(对BC的输出进行单步Q加权平均),并进行仅微调Q函数、保留BC权重不变的在线自改进。在LIBERO和双臂RoboTwin上,十次自改进迭代提升了所有测试基准分数(LIBERO-10从93%升至99%,RoboTwin从83.8%升至91.4%),并缩短了接近上限的套件(LIBERO-Object、LIBERO-Goal)的成功回合时长。在两个接触丰富的双臂真实机器人任务中,相同的循环(BC冻结,无人类干预)仅从自身部署轨迹中改进:叠杯子从40%升至90%,插入钱包从25%升至80%,迭代五次;而仅在成功轨迹上的SFT则停滞在55%和30%。在相同的在线预算下,Q规划是Best-of-N、过滤式SFT、IBRL、DSRL和DAWR中唯一能从失败中稳定改进且无需训练辅助演员的方法。

英文摘要

Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from that failure without additional human demonstrations. Reinforcement Learning fine-tuning offers a path to self-improvement but has proven difficult to scale to the multi-billion-parameter models underpinning modern robot policies. We propose Q-Planning, which equips a large visuomotor BC policy with a small off-policy Q-function. Because a Q-function estimates value rather than imitates actions, it can be trained on the same successful demonstrations as the BC policy and later absorb both successful and failed deployment rollouts, an asymmetry BC does not have. We exploit this asymmetry to enable value-guided action selection at inference (a single-step Q-weighted average over BC draws) and online self-improvement that fine-tunes only the Q-function, leaving the BC weights untouched. On LIBERO and bimanual RoboTwin, ten iterations of self-improvement lift every benchmark score we tested (LIBERO-10 93% to 99%, RoboTwin 83.8% to 91.4%) and shorten successful episodes on the near-ceiling suites (LIBERO-Object, LIBERO-Goal). On two contact-rich bimanual real-robot tasks, the same loop (BC frozen, no human intervention) improves purely from its own deployment rollouts: stack-cups 40% to 90% and insert-wallet 25% to 80% in five iterations, whereas SFT on successful rollouts alone stalls at 55% and 30%. Under an identical online budget Q-Planning is the only method, among Best-of-N, filtered SFT, IBRL, DSRL, and DAWR, that improves stably from failures without training an auxiliary actor.

CommentsProject page with videos: https://varungiridhar.github.io/qplanning/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑